Nemotron 3 Ultra
NVIDIA's largest open Nemotron 3 model, 550B parameters, for agentic reasoning.
Price
What the gateway charges per million tokens: what you send in, and what the model writes back.
Source: Vercel AI Gateway, checked 2026-09-25.
- Input
- $0.60per 1M tokens
- Output
- $2.40per 1M tokensOutput costs 4x the input.
Specs
How much it reads at once, how much it writes back, and how recent its knowledge is.
Sources: Vercel AI Gateway, checked 2026-09-25; Hugging Face, checked 2026-09-25.
- Context window
- 1M tokens1,000,000
- Max output
- 65K tokens65,000
- Knowledge cutoff
- Not listed
- Released
- Type
- Language
- Modalities
- Takes text. Returns text.
- License
- other
- Gateway ID
nvidia/nemotron-3-ultra-550b-a55b
Capabilities
What the gateway says this model supports.
Source: Vercel AI Gateway, checked 2026-09-25.
- Reasoning
- Tool use
- Structured output
- Implicit caching
Data handling
What happens to your prompts, and what you have to switch on.
Source: Vercel AI Gateway, checked 2026-09-25.
ZDREvery provider
Zero data retention
Available on a paid Vercel plan through the AI Gateway, from every provider on it (you turn it on).
NTEvery provider
No training on prompts
Offered by every provider on the Vercel AI Gateway, when you turn it on.
Usage
How much people actually run it, as counted where it is served.
- downloads
- 171.3K171,251 downloads
Counted on Hugging Face over the 30 days to 2026-09-25.
- requests
- 11.2M11,202,617 requests
- tokens
- 219.6B219,590,878,388 tokens
Counted on OpenRouter over the 31 days to 2026-09-25.
Friday
The brief, by email
Five to seven links an issue, each with what to do about it.
Unsubscribe from any issue.