NVIDIA API Pricing
NVIDIA publishes the open Nemotron family — models post-trained for reasoning, agentic workflows, and tool use, including Llama-Nemotron variants. They're designed to run efficiently on NVIDIA hardware and are freely available for self-hosting.
Models
4 tracked
All tiers, latest pricing.
All NVIDIA Models
| Model | Tier | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Nemotron 3 Nano 30B A3B | fast | $0.050 | $0.200 | 262K |
| Nemotron 3 Super | fast | $0.085 | $0.400 | 1M |
| Nemotron 3.5 Lightning | fast | $0.100 | $0.250 | 1M |
| Nemotron 3 Ultra | flagship | $0.600 | $3.60 | 512K |
Model Details
Nemotron 3 Nano 30B A3B
$0.050 inNVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems.
Nemotron 3 Super
$0.085 inNVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications.
Nemotron 3.5 Lightning
$0.100 inNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total.
Nemotron 3 Ultra
$0.600 inNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE).
Calculate NVIDIA API Costs
Use the TokenRate calculator to estimate exactly what NVIDIA models will cost for your workload.
Open Calculator →Other Providers