Llama 3.2 11B Vision Pricing
FastMeta · 128K tokens context
Llama 3.2 11B Vision from Meta costs $0.160 per 1 million input tokens and $0.160 per 1 million output tokens as of June 2026. The model supports a 128,000-token context window (approximately 96,000 words) with a 4K-token maximum output. A typical 1,000-token request costs $0.0002 in input charges; a 10,000-token request costs $0.0016.
| Input price | $0.160 / 1M tokens |
|---|---|
| Output price | $0.160 / 1M tokens |
| Output / input ratio | 1.0× |
| Context window | 128,000 tokens (~96,000 words) |
| Maximum output | 4,096 tokens |
| Cost per 1K tokens (input) | $0.0002 |
| Tier | Fast |
| Last verified |
Llama 3.2 11B Vision is Meta's small open-weight multimodal model — capable of understanding images at a fraction of GPT-4o's cost. The go-to for budget image + text pipelines.
Input Price
$0.160
per 1 million tokens
Output Price
$0.160
per 1 million tokens
Context Window
128K tokens
max 4K output
Cost Examples
| Request Type | Tokens | Input Cost | Output Cost |
|---|---|---|---|
| 1,000 word article | 1,333 | $0.000213 | $0.000064 |
| 10-page document (2,500 words) | 3,333 | $0.000533 | $0.00016 |
| 1,000 lines of code | 5,000 | $0.0008 | $0.00024 |
| 100K token document | 100,000 | $0.016 | $0.0048 |
Output cost estimated at 30% of input token count. Use the calculator for exact figures.
Strengths
- ✓Multimodal (image + text) at $0.16/1M
- ✓Open weights — self-hostable
- ✓Fast
Limitations
- –Below GPT-4o Vision on complex visual reasoning
- –Small output limit
Best Use Cases
Calculate Llama 3.2 11B Vision Costs
Use the TokenRate calculator to convert any budget, token count, or text into exact Llama 3.2 11B Vision costs — and compare across all models.
Open Calculator →Llama 3.2 11B Vision — FAQ
Related Models
Related Guides