TokenRate

Providers

AI API Providers

Each major AI provider with its pricing, model lineup, and where it shines.

Anthropic

19 models

Anthropic builds the Claude family of language models — known for strong reasoning, long context, and a safety-first design philosophy. Models are available via the Anthropic API and aggregators like AWS Bedrock and Google Vertex.

OpenAI

67 models

OpenAI created GPT and the modern LLM ecosystem. The GPT-4o and o-series models power ChatGPT, the OpenAI API, and Microsoft Azure OpenAI.

Google

30 models

Google's Gemini family pushes the long-context frontier with 1M+ token windows, native multimodality, and aggressive pricing on the Flash tier. Available via Google AI Studio and Vertex AI.

Meta

11 models

Meta publishes the Llama family of open-weight models. Llama 3.1 ranges from a tiny 8B variant to a 405B frontier model, and is hosted by every major inference provider.

DeepSeek

11 models

DeepSeek is a Chinese AI lab whose open-weight V3 and R1 models redefined the cost frontier for reasoning and general capability.

xAI

8 models

xAI builds the Grok family of models, integrated with real-time X (Twitter) data. Available through the xAI API and X's premium tier.

Mistral

22 models

Mistral AI is a French AI lab building open and proprietary models with a focus on European languages, function calling, and code-specialized variants.

Qwen

47 models

Qwen is Alibaba's family of open-weight language and multimodal models. The Qwen 2.5 and QwQ series consistently top open-source benchmarks for coding and math, with strong multilingual support across Asian languages.

Cohere

6 models

Cohere builds enterprise-focused LLMs optimized for retrieval-augmented generation (RAG) and tool use. The Command R series is purpose-built for grounded, citation-accurate answers in document Q&A pipelines.

Amazon

5 models

Amazon's Nova family of generative AI models is natively integrated with AWS Bedrock, offering text, image, and video understanding across Pro, Lite, and Micro tiers — optimized for enterprise AWS workloads.

Microsoft

2 models

Microsoft's Phi family of small language models (SLMs) is designed for edge and on-device inference. Phi-4 delivers frontier-class reasoning within a 14B parameter footprint, optimized for STEM and coding tasks.

Zhipu AI

12 models

Zhipu AI (Z.ai) is a Beijing-based lab spun out of Tsinghua University, building the open-weight GLM series. The GLM-4.5, 4.6, 4.7 and GLM-5 models are strong agentic and coding performers with vision (GLM-V) variants, popular for self-hosting at a fraction of frontier API prices.

Moonshot AI

7 models

Moonshot AI is the Chinese lab behind the Kimi family. Kimi K2 is a large open-weight Mixture-of-Experts model tuned for agentic tool use, coding, and very long context — a leading open alternative to closed frontier models at low API cost.

NVIDIA

3 models

NVIDIA publishes the open Nemotron family — models post-trained for reasoning, agentic workflows, and tool use, including Llama-Nemotron variants. They're designed to run efficiently on NVIDIA hardware and are freely available for self-hosting.

MiniMax

8 models

MiniMax is a Shanghai-based lab whose MiniMax-M series pairs very long context windows (up to 1M+ tokens) with strong agentic and coding performance at aggressive prices. The models are open-weight and widely hosted.

Perplexity

5 models

Perplexity builds the Sonar family — models with built-in web search that return grounded, citation-backed answers. Sonar, Sonar Pro, and the Sonar Reasoning and Deep Research tiers are the engine behind Perplexity's answer product.

Nous Research

4 models

Nous Research builds the open-weight Hermes series — steerable, instruction-tuned fine-tunes of Llama with strong function-calling and minimal refusals. A favourite of the open-source community for agents and self-hosting.

ByteDance

5 models

ByteDance (TikTok's parent) ships the Seed family of general-purpose models and UI-TARS, a GUI-agent model that controls computers from screenshots. The Seed models offer long context and competitive pricing across flagship and flash tiers.

Arcee AI

2 models

Arcee AI builds small, efficient enterprise models — the Virtuoso, Coder, and Trinity families — using model-merging and distillation. They target strong quality-per-dollar for coding, reasoning, and on-prem deployment.

AI21 Labs

1 models

AI21 Labs is an Israeli lab whose Jamba models use a hybrid Mamba-Transformer architecture for efficient very-long-context inference. Jamba targets enterprise document and RAG workloads with 256K-token windows.

Reka AI

2 models

Reka AI is a research lab building compact multimodal models — the Reka Flash and Edge series understand text, images, audio, and video. Designed to be efficient enough for on-device and edge deployment.

IBM

2 models

IBM's open-weight Granite family targets enterprise use — small, efficient, commercially-licensed models with strong governance and transparency. Granite is tuned for coding, RAG, and tool use within IBM watsonx.

Tencent

3 models

Tencent builds the Hunyuan family of models, including Mixture-of-Experts variants. Hunyuan powers Tencent's products and is available open-weight for general chat, reasoning, and multilingual workloads.

Inflection AI

2 models

Inflection AI builds emotionally-intelligent models behind Pi, its personal AI. The Inflection-3 models — Pi (conversational) and Productivity (instruction-following) — target empathetic, safe enterprise assistants.

Liquid AI

0 models

Liquid AI, an MIT spin-out, builds Liquid Foundation Models (LFMs) using a non-transformer architecture for high efficiency. LFMs are tiny, fast, and designed for on-device and edge inference with low memory footprints.

Allen Institute for AI

1 models

The Allen Institute for AI (Ai2) builds OLMo — fully open models that release weights, training data, and code together. OLMo is a reference point for reproducible, transparent open-source AI research.

Baidu

1 models

Baidu builds the ERNIE family, China's longest-running large-model line. ERNIE 4.5 includes large Mixture-of-Experts and vision-language variants, available open-weight and via Baidu's cloud.

Writer

1 models

Writer builds the Palmyra family of enterprise models tuned for business writing, domain-specific knowledge work, and long-context document tasks. Palmyra X5 offers a 1M-token window for whole-document workflows.

Upstage

1 models

Upstage is a South Korean lab building the Solar family — small, efficient models that punch above their parameter count via depth-up-scaling. Solar Pro targets strong document understanding and multilingual performance at low cost.