Meta API Pricing
Meta publishes the Llama family of open-weight models. Llama 3.1 ranges from a tiny 8B variant to a 405B frontier model, and is hosted by every major inference provider.
Models
11 tracked
All tiers, latest pricing.
All Meta Models
| Model | Tier | Input / 1M | Output / 1M | Context |
|---|---|---|---|---|
| Llama 3.2 1B | fast | $0.027 | $0.201 | 60K |
| Llama 3.1 8B | fast | $0.050 | $0.080 | 131K |
| Llama 3.2 3B | fast | $0.051 | $0.335 | 131K |
| Llama 4 Scout | fast | $0.100 | $0.300 | 1M |
| Llama 3.3 70B | balanced | $0.130 | $0.400 | 131K |
| Llama 3.2 11B Vision | fast | $0.160 | $0.160 | 128K |
| Llama Guard 4 12B | fast | $0.180 | $0.180 | 1M |
| Llama 4 Maverick | balanced | $0.200 | $0.800 | 1M |
| Llama 3.1 70B | balanced | $0.400 | $0.400 | 131K |
| Llama 3.2 90B Vision | balanced | $0.900 | $0.900 | 128K |
| Llama 3.1 405B | flagship | $2.70 | $2.70 | 128K |
Model Details
Llama 3.2 1B
$0.027 inLlama 3.2 1B is one of the smallest capable LLMs — sub-$0.05 per million tokens and runs on CPUs. Useful for classification and routing at extreme scale or very constrained hardware.
Llama 3.1 8B
$0.050 inLlama 3.1 8B is the smallest open-weight Llama 3.1 model — extremely cheap to host or call, and good enough for classification, extraction, and basic chat.
Llama 3.2 3B
$0.051 inLlama 3.2 3B is a tiny but surprisingly capable open-weight model — one of the cheapest LLMs available from any provider. Fits on edge hardware and consumer GPUs with room to spare.
Llama 4 Scout
$0.100 inLlama 4 Scout is Meta's latest MoE model with an industry-leading 10M token context window at an affordable price. Remarkable context-to-cost ratio — suitable for entire-codebase and very long document tasks.
Llama 3.3 70B
$0.130 inLlama 3.3 70B improves on Llama 3.1 70B with better instruction-following and reasoning, at the same price point. The recommended Llama 70B for new projects — same hosting cost, meaningfully better quality.
Llama 3.2 11B Vision
$0.160 inLlama 3.2 11B Vision is Meta's small open-weight multimodal model — capable of understanding images at a fraction of GPT-4o's cost. The go-to for budget image + text pipelines.
Llama Guard 4 12B
$0.180 inLlama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification.
Llama 4 Maverick
$0.200 inLlama 4 Maverick is Meta's flagship open-weight model in the Llama 4 generation — multimodal, 1M context, and competitive with GPT-4o at a fraction of the API cost.
Llama 3.1 70B
$0.400 inLlama 3.1 70B is the mid-size open-weight model in the 3.1 family — a popular sweet spot for production workloads that need GPT-4o-mini-class quality at open-weight prices.
Llama 3.2 90B Vision
$0.900 inLlama 3.2 90B Vision is Meta's large open-weight multimodal model — strong on image understanding and visual reasoning while remaining self-hostable. Best open multimodal option before Llama 4.
Llama 3.1 405B
$2.70 inLlama 3.1 405B is Meta's largest open-weight model — competitive with GPT-4-class models on many benchmarks and uniquely available for self-hosting. Symmetric input/output pricing is common across hosted providers.
Calculate Meta API Costs
Use the TokenRate calculator to estimate exactly what Meta models will cost for your workload.
Open Calculator →Other Providers