A large language model customized by NVIDIA to improve the helpfulness of LLM generated responses. It is a fine-tuned version of Llama 3.1 70B Instruct. The model was trained using RLHF (REINFORCE) with HelpSteer2-Preference prompts.
Organization
NVIDIA
GPU and AI company
旗下模型
Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.
Safety model for policy screening, moderation, and risk-aware routing workflows
Flagship Nemotron model for high-throughput reasoning and complex agents
A 253B parameter derivative of Meta Llama 3.1 405B Instruct, developed by NVIDIA using Neural Architecture Search (NAS) and vertical compression. It underwent multi-phase post-training (SFT for Math, Code, Reasoning, Chat, Tool Calling; RL with GRPO) to enhance reasoning and instruction-following. Optimized for accuracy/efficiency tradeoff on NVIDIA GPUs. Supports 128k context.
Nemotron model for efficient reasoning, coding, and specialized AI agents
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Reranking model for improving retrieval quality in search and recommendation systems
Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) derived from Meta Llama-3.3-70B-Instruct. It's post-trained for reasoning, chat, RAG, and tool calling, offering a balance between accuracy and efficiency (optimized for single H100). It underwent multi-phase post-training including SFT and RL (RLOO, RPO).
Mistral Nemotron
NVIDIAMistral model for multilingual chat, reasoning, and tool-assisted workflows
Nemotron 3 Content Safety
NVIDIASafety model for policy screening, moderation, and risk-aware routing workflows
Nemotron 3 Nano 30B A3B
NVIDIASmall Nemotron 3 MoE for efficient coding, math, and long-context agents
Open Nemotron omni model combining reasoning with text, vision, and audio
Nemotron 3 Super 120B A12B
NVIDIANemotron middle tier for collaborative agents and high-volume reasoning workloads
Nemotron 3 Ultra 550B A55B
NVIDIALargest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
Nemotron 3.5 Content Safety
NVIDIASafety model for policy screening, moderation, and risk-aware routing workflows
Nemotron Cascade 2 30B A3B
NVIDIANemotron model for efficient reasoning, coding, and specialized AI agents
Safety model for policy screening, moderation, and risk-aware routing workflows
Nemotron Mini 4B Instruct
NVIDIACompact Nemotron model for efficient reasoning and deployable AI agents
Nemotron Nano 12B v2 VL
NVIDIANemotron multimodal model for visual reasoning and agentic AI workflows
Nemotron Nano 9B v2
NVIDIANVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.
Nemotron VoiceChat
NVIDIANemotron multimodal model for visual reasoning and agentic AI workflows