TokenRate

Meta API Pricing

Meta publishes the Llama family of open-weight models. Llama 3.1 ranges from a tiny 8B variant to a 405B frontier model, and is hosted by every major inference provider.

Official site: llama.meta.com

Cheapest

Llama 3.2 1B

$0.027/1M in

Flagship

Llama 3.1 405B

$2.70/1M in

Models

11 tracked

All tiers, latest pricing.

All Meta Models

ModelTierInput / 1MOutput / 1MContext
Llama 3.2 1Bfast$0.027$0.20160K
Llama 3.1 8Bfast$0.050$0.080131K
Llama 3.2 3Bfast$0.051$0.335131K
Llama 4 Scoutfast$0.100$0.3001M
Llama 3.3 70Bbalanced$0.130$0.400131K
Llama 3.2 11B Visionfast$0.160$0.160128K
Llama Guard 4 12Bfast$0.180$0.1801M
Llama 4 Maverickbalanced$0.200$0.8001M
Llama 3.1 70Bbalanced$0.400$0.400131K
Llama 3.2 90B Visionbalanced$0.900$0.900128K
Llama 3.1 405Bflagship$2.70$2.70128K

Model Details

Llama 3.2 1B

$0.027 in

Llama 3.2 1B is one of the smallest capable LLMs — sub-$0.05 per million tokens and runs on CPUs. Useful for classification and routing at extreme scale or very constrained hardware.

Llama 3.1 8B

$0.050 in

Llama 3.1 8B is the smallest open-weight Llama 3.1 model — extremely cheap to host or call, and good enough for classification, extraction, and basic chat.

Llama 3.2 3B

$0.051 in

Llama 3.2 3B is a tiny but surprisingly capable open-weight model — one of the cheapest LLMs available from any provider. Fits on edge hardware and consumer GPUs with room to spare.

Llama 4 Scout

$0.100 in

Llama 4 Scout is Meta's latest MoE model with an industry-leading 10M token context window at an affordable price. Remarkable context-to-cost ratio — suitable for entire-codebase and very long document tasks.

Llama 3.3 70B

$0.130 in

Llama 3.3 70B improves on Llama 3.1 70B with better instruction-following and reasoning, at the same price point. The recommended Llama 70B for new projects — same hosting cost, meaningfully better quality.

Llama 3.2 11B Vision

$0.160 in

Llama 3.2 11B Vision is Meta's small open-weight multimodal model — capable of understanding images at a fraction of GPT-4o's cost. The go-to for budget image + text pipelines.

Llama Guard 4 12B

$0.180 in

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification.

Llama 4 Maverick

$0.200 in

Llama 4 Maverick is Meta's flagship open-weight model in the Llama 4 generation — multimodal, 1M context, and competitive with GPT-4o at a fraction of the API cost.

Llama 3.1 70B

$0.400 in

Llama 3.1 70B is the mid-size open-weight model in the 3.1 family — a popular sweet spot for production workloads that need GPT-4o-mini-class quality at open-weight prices.

Llama 3.2 90B Vision

$0.900 in

Llama 3.2 90B Vision is Meta's large open-weight multimodal model — strong on image understanding and visual reasoning while remaining self-hostable. Best open multimodal option before Llama 4.

Llama 3.1 405B

$2.70 in

Llama 3.1 405B is Meta's largest open-weight model — competitive with GPT-4-class models on many benchmarks and uniquely available for self-hosting. Symmetric input/output pricing is common across hosted providers.

Calculate Meta API Costs

Use the TokenRate calculator to estimate exactly what Meta models will cost for your workload.

Open Calculator →

Other Providers