Organization

NVIDIA

US 22 个模型

GPU and AI company

旗下模型

开源

A large language model customized by NVIDIA to improve the helpfulness of LLM generated responses. It is a fine-tuned version of Llama 3.1 70B Instruct. The model was trained using RLHF (REINFORCE) with HelpSteer2-Preference prompts.

2024年10月1日
开源

Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.

2025年3月18日
开源

Flagship Nemotron model for high-throughput reasoning and complex agents

2025年4月7日 128K 上下文 $0.60 起
开源

A 253B parameter derivative of Meta Llama 3.1 405B Instruct, developed by NVIDIA using Neural Architecture Search (NAS) and vertical compression. It underwent multi-phase post-training (SFT for Math, Code, Reasoning, Chat, Tool Calling; RL with GRPO) to enhance reasoning and instruction-following. Optimized for accuracy/efficiency tradeoff on NVIDIA GPUs. Supports 128k context.

2025年4月7日
开源 多模态

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

2026年2月10日
开源 多模态

Reranking model for improving retrieval quality in search and recommendation systems

2026年3月31日
开源

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) derived from Meta Llama-3.3-70B-Instruct. It's post-trained for reasoning, chat, RAG, and tool calling, offering a balance between accuracy and efficiency (optimized for single H100). It underwent multi-phase post-training including SFT and RL (RLOO, RPO).

2025年3月18日
开源
N

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

2025年6月11日
开源

Safety model for policy screening, moderation, and risk-aware routing workflows

2026年4月16日
开源

Small Nemotron 3 MoE for efficient coding, math, and long-context agents

2025年12月15日 262K 上下文 $0.00 起
开源

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

2026年3月11日 262K 上下文 $0.00 起
开源

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

2026年6月4日 1.0M 上下文 $0.00 起
开源 多模态

Safety model for policy screening, moderation, and risk-aware routing workflows

2026年6月4日
开源

Nemotron model for efficient reasoning, coding, and specialized AI agents

2026年3月24日
开源

Compact Nemotron model for efficient reasoning and deployable AI agents

2024年8月21日
开源 多模态

Nemotron multimodal model for visual reasoning and agentic AI workflows

2025年10月28日 128K 上下文 $0.20 起
开源

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the model to generate reasoning traces first generally results in higher-quality final solutions to queries and tasks.

2025年8月18日
开源 多模态

Nemotron multimodal model for visual reasoning and agentic AI workflows

2026年3月16日