TokenRate

NVIDIA API Pricing

NVIDIA publishes the open Nemotron family — models post-trained for reasoning, agentic workflows, and tool use, including Llama-Nemotron variants. They're designed to run efficiently on NVIDIA hardware and are freely available for self-hosting.

Official site: build.nvidia.com

Cheapest

Nemotron 3 Nano 30B A3B

$0.050/1M in

Flagship

Nemotron 3 Ultra

$0.600/1M in

Models

4 tracked

All tiers, latest pricing.

All NVIDIA Models

ModelTierInput / 1MOutput / 1MContext
Nemotron 3 Nano 30B A3Bfast$0.050$0.200262K
Nemotron 3 Superfast$0.085$0.4001M
Nemotron 3.5 Lightningfast$0.100$0.2501M
Nemotron 3 Ultraflagship$0.600$3.60512K

Model Details

Calculate NVIDIA API Costs

Use the TokenRate calculator to estimate exactly what NVIDIA models will cost for your workload.

Open Calculator →

Other Providers