Largest Context Window LLM Models

This page surfaces models with the biggest context windows for retrieval, long-document analysis, codebase review, and agent workflows.

50Models listed
1M input + 500K outputCost example tokens
USD / 1MNormalized prices

Quick shortlist

Start with Grok 4.20 Multi-Agent.

This guide is sorted by context window, so the first rows are the strongest starting point for RAG, long documents, and large codebase context.

Lead model Grok 4.20 Multi-Agent
ProviderSpaceXAI
Sample cost$2.5
Context2M

The ranking is a discovery aid, not a final recommendation. Always compare the model against your workload and verify provider pricing before production use.

How to read this ranking

Models are sorted by context window size. Use this page when your workflow needs long documents, large retrieval payloads, or multi-file context.

Model Ranking

Browse all models
ModelProviderPromptOutputSample costYour CostContextPopularityRelease
Grok 4.20 Multi-AgentSpaceXAI$1.25$2.5$2.5$2.52M
Grok 4.20SpaceXAI$1.25$2.5$2.5$2.52M
🔥DeepSeek V4 Flash 0731DeepSeek$0.08$0.18$0.17$0.171.31M#1
DeepSeek V4 Flash Latestdeepseek$0.04$0.13$0.11$0.111.31M
Llama 4 ScoutMeta$0.1$0.3$0.25$0.251.31M
🔥MiMo-V2.5Xiaomi$0.14$0.28$0.28$0.281.05M#2
🔥GPT-5.6 LunaOpenAI$0.2$1.2$0.8$0.81.05M#6
🔥GPT-5.6 SolOpenAI$2$10$7$71.05M#15
GPT-5.6 Luna ProOpenAI$0.2$1.2$0.8$0.81.05M
GPT-5.6 Luna Pro (batch)OpenAI$0.1$0.6$0.4$0.41.05M
GPT-5.6 Luna (batch)OpenAI$0.1$0.6$0.4$0.41.05M
GPT-5.6 Terra ProOpenAI$2$12$8$81.05M
GPT-5.6 Terra Pro (batch)OpenAI$1$6$4$41.05M
GPT-5.6 TerraOpenAI$2$12$8$81.05M
GPT-5.6 Terra (batch)OpenAI$1$6$4$41.05M
GPT-5.6 Sol ProOpenAI$2$10$7$71.05M
GPT-5.6 Sol Pro (batch)OpenAI$1$5$3.5$3.51.05M
GPT-5.6 Sol (batch)OpenAI$1$5$3.5$3.51.05M
OpenAI GPT LatestOpenAI$2$10$7$71.05M
GPT-5.5 ProOpenAI$30$180$120$1201.05M
GPT-5.5 Pro (batch)OpenAI$15$90$60$601.05M
GPT-5.5OpenAI$5$30$20$201.05M
GPT-5.5 (batch)OpenAI$2.5$15$10$101.05M
MiMo-V2.5-ProXiaomi$0.435$0.87$0.87$0.871.05M
GPT-5.4 ProOpenAI$30$180$120$1201.05M
GPT-5.4 Pro (batch)OpenAI$15$90$60$601.05M
GPT-5.4OpenAI$2.5$15$10$101.05M
GPT-5.4 (batch)OpenAI$1.25$7.5$5$51.05M
LongCat 2.0Meituan$0.3$1.2$0.9$0.91.05M
Owl AlphaOpenRouter$0$0$0$01.05M
New🔥Ox Alphastealth$0$0$0$01.05M#4
🔥DeepSeek V4 Flash 0423DeepSeek$0.0517$0.1033$0.1$0.11.05M#5
🔥GLM 5.2Z.ai$0.966$3.036$2.48$2.481.05M#8
🔥DeepSeek V4 Pro 0423DeepSeek$0.3969$0.7938$0.79$0.791.05M#10
New🔥Gemini 3.7 FlashGoogle$0.375$1.875$1.31$1.311.05M#12
🔥Gemini 3.6 FlashGoogle$0.75$3.75$2.62$2.621.05M#13
🔥Kimi K3MoonshotAI$3$15$10.5$10.51.05M#14
🔥Gemini 3 Flash PreviewGoogle$0.5$3$2$21.05M#18
NewMuse Spark 1.2 ContributorMeta$0.1$0.2$0.2$0.21.05M
NewDeepSeek V4 Flash Vision ExpDeepSeek$0.22$0.66$0.55$0.551.05M
NewGLM LatestZ.ai$1.4$4.4$3.6$3.61.05M
NewGLM 5.3Z.ai$1.4$4.4$3.6$3.61.05M
NewGemini 3.7 Flash (batch)Google$0.1875$0.9375$0.66$0.661.05M
NewQwen3.8 2.4T A95BQwen$2$6$5$51.05M
NewDeepSeek V4 Pro 0813DeepSeek$1.122$3.366$2.81$2.811.05M
Muse Spark 1.2Meta$1.25$4.25$3.38$3.381.05M
Inkling SmallThinking Machines$0.45$1.2$1.05$1.051.05M
Laguna S 2.1Poolside$0.09$0.18$0.18$0.181.05M
Gemini 3.6 Flash (batch)Google$0.375$1.875$1.31$1.311.05M
Gemini 3.5 Flash LiteGoogle$0.3$2.5$1.55$1.551.05M

Pricing FAQ

How is the sample workload cost calculated?

The sample workload uses 1,000,000 input tokens plus 500,000 output tokens, then applies each model's normalized USD price per 1 million tokens.

Why do input and output token prices matter separately?

Many applications are output-token heavy, while retrieval and classification workloads may be input-token heavy. Comparing both prices helps avoid picking a model that is cheap for the wrong workload shape.

Should I verify prices before production use?

Yes. AI Model Matrix normalizes public pricing metadata for comparison, but provider availability, limits, and prices can change. Always verify the final contract or provider dashboard before production use.

Related Guides

Cheapest LLM APIs

Sort models by estimated workload cost and normalized token prices.

Open guide

Largest Context Windows

Find models for long documents, retrieval, and codebase context.

Open guide

Coding Models

Compare code-oriented models by cost, context, and practical popularity signals.

Open guide

Free Models

Browse zero-price models for prototypes and evaluation.

Open guide

RAG Models

Start from large context windows and practical input-cost constraints.

Open guide

Chatbot Costs

Find budget-sensitive models for output-heavy assistant traffic.

Open guide

Cost Calculator

Enter your own input and output token volume before narrowing the shortlist.

Estimate cost