Nemotron-3-120B-Super
Modelfamilie van NVIDIA · gebruikt door 1 applicatie
Nemotron-3-120B-Super is a large language model (LLM) developed by NVIDIA, featuring 120 billion parameters. It is part of the Nemotron-3 family of models designed for high-performance text generation, synthetic data generation, and enterprise applications. The model is optimized for use in data centers and cloud environments, leveraging NVIDIA's advanced GPU infrastructure. It supports a wide range of natural language processing (NLP) tasks, including but not limited to text generation, summarization, translation, and code generation. Nemotron-3-120B-Super is particularly notable for its use in generating synthetic data to train other AI models, which can help address data scarcity and privacy concerns in AI development. NVIDIA Nemotron-3-Super-120B-A12B-BF16 is a 120B-parameter (12B active) Latent Mixture-of-Experts (LatentMoE) hybrid model (Mamba-2 + MoE + Attention) with Multi-Token Prediction (MTP). It is optimized for agentic workflows, long-context reasoning (up to 1M tokens), tool use, and RAG, and is commercially usable under the NVIDIA Nemotron Open Model License. The model excels in math, code, science, and multilingual tasks, and is designed for high-volume workloads like IT ticket automation.
Zes dimensies, gekoppeld aan bewijs
Werk je bij NVIDIA? Claim deze vermelding om de gegevens te corrigeren of aan te vullen.