Skip to content

Nemotron 3 Ultra

NVIDIA's largest open Nemotron 3 model, 550B parameters, for agentic reasoning.

By NVIDIAsource

Facts read from official pages, last on

Official page

Price

What the gateway charges per million tokens: what you send in, and what the model writes back.

Source: Vercel AI Gateway, checked 2026-09-25.

Input
$0.60per 1M tokens
Output
$2.40per 1M tokensOutput costs 4x the input.

Specs

How much it reads at once, how much it writes back, and how recent its knowledge is.

Sources: Vercel AI Gateway, checked 2026-09-25; Hugging Face, checked 2026-09-25.

Context window
1M tokens1,000,000
Max output
65K tokens65,000
Knowledge cutoff
Not listed
Released
Type
Language
Modalities
Takes text. Returns text.
License
other
Gateway ID
nvidia/nemotron-3-ultra-550b-a55b

Capabilities

What the gateway says this model supports.

Source: Vercel AI Gateway, checked 2026-09-25.

  • Reasoning
  • Tool use
  • Structured output
  • Implicit caching

Data handling

What happens to your prompts, and what you have to switch on.

Source: Vercel AI Gateway, checked 2026-09-25.

ZDREvery provider

Zero data retention

Available on a paid Vercel plan through the AI Gateway, from every provider on it (you turn it on).

NTEvery provider

No training on prompts

Offered by every provider on the Vercel AI Gateway, when you turn it on.

Usage

How much people actually run it, as counted where it is served.

downloads
171.3K171,251 downloads

Counted on Hugging Face over the 30 days to 2026-09-25.

requests
11.2M11,202,617 requests
tokens
219.6B219,590,878,388 tokens

Counted on OpenRouter over the 31 days to 2026-09-25.

Friday

The brief, by email

Five to seven links an issue, each with what to do about it.

Unsubscribe from any issue.

This page as Markdown