Table of Contents
- 1. 1M Support Is Everywhere — But "Reads to the End" Is Another Matter
- 2. What Is Context? — Separate the Container from Its Contents
- 3. Read the Container Size from Three Numbers
- 4. Three Reasons "Bigger Is Better" Doesn't Hold
- 5. The Cost Trap — Models Whose Rates Jump on Long Inputs, and Models Whose Rates Don't
- 6. Five Saving Tactics — Ranked by Real Impact for Solo Devs
- Summary
- FAQ
In 2023, a 32K-token context window felt "spacious." Today, a window in the 1-million-token (1M) class is simply assumed for top-tier models. As of September 2026, Anthropic, OpenAI, and Google all list an input limit of roughly one million tokens in the official specs of their top models. Individual model names and figures change with every release, so I leave those to the list of current models and cutoff dates and to each vendor's official pages. This article focuses on what survives a generation change: how to read the numbers, and how to work with them.
"One million tokens" translates to roughly 8–10 paperback books in English, or tens of thousands of lines of source code. We can now keep that much "in view" within a single session. But a document fitting in the container is not the same as the model reading all of it. The scores OpenAI publishes on its multi-needle long-context benchmark (MRCR) show that the same model scores lower as the input gets longer, and how far it drops varies widely by model and generation (covered in §1 and §4).
Let me get my take out front: the era of choosing a model purely on container size is over. What matters is the trio of "effective context × cost × how you feed it." This article walks through what context actually is, how to read a spec sheet, why bigger isn't enough, how to measure the effective range for your own use case, how long-input pricing jumps, and five saving tactics that solo developers and small teams can apply today — using official figures and published benchmark numbers.
The Container Grew 250x in Three Years
— A timeline of how 1M went from luxury to baseline
But "supported" and "read to the end" are different things. On OpenAI's long-context benchmark MRCR (8 needles), GPT-5.5 goes from 98.1% at 4K–8K to 74.0% at 512K–1M.
How far scores fall differs by model and generation (evaluation table in OpenAI's "Introducing GPT-5.5," April 23, 2026; details in §1 and §4).
1. 1M Support Is Everywhere — But "Reads to the End" Is Another Matter
1M support spread fast over the past two years. In April 2025 OpenAI's GPT-4.1 shipped with about 1.05 million tokens, in August 2025 Anthropic's Claude Sonnet 4 reached 1 million tokens (in beta), and as of September 2026 the top models from Anthropic, OpenAI, and Google all list roughly one million tokens in their official specs. Back in 2023, 32K felt spacious — that's more than 30x in just three years. The container-size race looked finished.
But look at the long-context scores the vendors publish themselves, and the picture gets more complicated. The benchmark with a full set of results by length is OpenAI's MRCR v2 (8 needles). It hides eight requests of the same kind — say, "write a poem about tapirs" — inside a long conversation with an AI, then asks for one specific instance, such as "return the 2nd poem," exactly. Because the model has to tell near-identical items apart, down to their order, it is a multi-needle needle-in-a-haystack test. The score is how closely the returned text matches the correct one. Here are the results by length:
- GPT-5.5: 98.1% at 4K–8K, 87.5% at 128K–256K, 74.0% at 512K–1M
- GPT-5.4 (the previous generation in the same table): 97.3% → 79.3% → 36.6% across the same three bands
- Claude Opus 4.7 (figures OpenAI listed in the same table): 59.2% at 128K–256K, 32.2% at 512K–1M
- GPT-6 Astra (announced September 2026): 100.0% at 256K–512K, 96.3% at 512K–1M (GPT-5.6 Sol in the same table: 91.5% and 73.8%)
Sources (checked September 26, 2026): evaluation tables in OpenAI's "Introducing GPT-5.5" (April 23, 2026) and "GPT-6 Astra" (September 3, 2026). How the benchmark works: the OpenAI MRCR dataset card. Model names are those at the time of each announcement.
Two things stand out. First, the same model scores lower as the input gets longer. Second, how steeply it drops varies widely by model and generation — in the band closest to 1M, GPT-5.4 fell below 40%, while GPT-6 Astra, announced in September 2026, stayed above 90%. Rankings reshuffle with every generation, so these numbers go stale quickly. What lasts is the lesson: the advertised limit and the range where accuracy actually holds are two different numbers.
Don't misread this. It's not "Claude or GPT is bad." Use cases that truly need all 1M are rarer than you'd think. If a model reads stably up to 300K (roughly 2–3 books), almost every coding, research, and summarization task fits. The problem is choosing on the "1M supported" number alone — that's how you end up with the wrong decision criteria.
2. What Is Context? — Separate the Container from Its Contents
Quick terminology. Three words get mixed up in this space.
Token, Window, Context
In short: "window = container size," "context = contents," "token = unit."
A big container with messy contents still gives you messy answers.
Also: don't confuse "context" with "memory." Context lives inside the session — close the chat and it's gone. Features like ChatGPT Memory or Claude Memory, on the other hand, are a separate cross-session retention mechanism. Memory contents do eventually get injected into the context window, but from the user's perspective it's persistent storage vs. ephemeral workspace.
3. Read the Container Size from Three Numbers
When you look at a model's spec sheet, there are only three things to check around context. Get these down and you can compare any model with the same procedure.
The number the marketing leads with. But it tells you how much fits in, not how much actually gets read — as the next section covers, the effective figure is far below it.
Often overlooked, but it is an order of magnitude smaller than the input limit. Even among major models as of September 2026, you can send in 1 million tokens but get back only about 65K–128K. For "rewrite this whole long document" jobs, this is the limit you hit first.
Flat across the whole window, or a threshold where the unit price jumps. This matters most in production and is usually the one thing missing from the spec sheet. §5 works the numbers out.
Below are the figures on each vendor's official pages as of September 29, 2026. The numbers move with each generation, so read this as a worked example of how the three axes above turn into real differences. For today's specific model names, see the list of current models and cutoff dates; for the numbers, check each vendor's official pages.
| Family (examples as of Sep 2026) | Input limit | Output limit | Long-input pricing |
|---|---|---|---|
| Anthropic top tier (Claude Fable 5.1, Opus 5.5, Sonnet 5.5) | 1,000,000 | 128,000 | Same rate all the way to the limit |
| Anthropic lightweight (Claude Haiku 4.5) | 200,000 | 64,000 | — |
| OpenAI (GPT-6 Astra, Sol) | 1,050,000 | 128,000 | Inputs over 272K: the whole request at 2x input, 1.5x output |
| Google (Gemini 3.1 Pro, preview) | 1,048,576 | 65,536 | Over 200K: input $2→$4, output $12→$18 |
Sources (checked September 29, 2026): Anthropic Models overview and Pricing / OpenAI model pages for GPT-6 Astra and GPT-6 Sol / Google Gemini 3.1 Pro Preview and Gemini API pricing
Look at the table and the input limits are nearly identical across vendors; the differences are in output limits and the shape of the pricing. When the limits match, window size stops being a reason to choose. What does differ: Anthropic keeps the same rate up to the limit, while OpenAI and Google raise the rate past a certain length — so under the same "1M support" label, how freely you can send long inputs varies. That's not just a pricing detail; it reflects different approaches to long-context workloads. The cost chapter runs the numbers.
In practice, the choice goes like this. Decide by the document sizes you routinely work with. If they fit under ~200K, choose on accuracy stability in that band and whether you stay below a pricing threshold, not on window size. Only if you routinely handle giant documents over 300K do the width of the effective range and the long-input rate become deciding factors. Don't commit to one model — split by use case. That way of deciding holds no matter which generation is current.
Where to check the limits
Check the numbers on each vendor's official pages, not in roundup sites or article tables (including this one). Where to look is fairly consistent:
- Anthropic: the comparison table on the models overview page has "Context window" and "Max output" rows. Long-input pricing is in the "Long context pricing" section of the pricing page
- OpenAI: each model's page in the API docs lists "context window" and "max output tokens," plus a note on pricing for long prompts
- Google: the Gemini API model page lists "Input token limit" and "Output token limit," and the pricing page has a "prompts > 200k tokens" tier
One more thing that's easy to miss: the same limit holds different amounts of text depending on the model. Limits are counted in tokens, so when the tokenizer (the scheme that splits text into tokens) changes, the token count of the same document changes too (see the note in §5). When you switch models, recount with your own documents to confirm you still have headroom.
4. Three Reasons "Bigger Is Better" Doesn't Hold
The spec table in the last chapter only shows the physical size of the container. So does the model actually use the container it advertises? Bottom line: don't assume it reads with the same accuracy all the way to the limit. There are three reasons.
Reason ①: Lost in the Middle
This is the phenomenon researchers from Stanford and elsewhere (Liu et al.) reported in 2023 in the paper "Lost in the Middle." On tasks such as finding an answer across multiple documents, models answered best when the answer sat at the beginning or end of the input and lost significant accuracy when it sat in the middle — even models built for long context showed the same pattern.
In practice it feels like this: "Paste a long PDF in full, ask 'What's the figure for X?', and it mixes up a number right around the middle." That's Lost in the Middle. How severe it is varies by model, but the safe assumption is that information placed mid-document is easier to miss — and to adjust how you feed it accordingly.
Reason ②: Context Rot
The longer a conversation runs, the more your early instructions fade. You asked for a formal tone at the start, and 20 turns later the model has drifted back to casual — that's Context Rot.
Two causes. ① Early instructions get treated as relatively old and light within the conversation history. ② The longer history spreads attention thin, making specific tokens harder to reference. In its developer documentation, Anthropic calls the drop in accuracy and recall as token counts grow "context rot," and writes that choosing what goes into context matters as much as how much fits. In September 2025 it laid out how to handle the problem as a deliberate skill in an engineering post titled "Effective context engineering for AI agents".
Reason ③: Advertised Context ≠ Effective Context
Take only the band closest to 1M (512K–1M) from the §1 figures and line them up. All come from OpenAI's announcement evaluation tables and use the same benchmark, OpenAI MRCR v2 (8 needles).
Near 1M, can the model return the one item asked for?
Sources: evaluation tables in OpenAI's "Introducing GPT-5.5" (April 23, 2026; GPT-5.5, GPT-5.4, Claude Opus 4.7) and "GPT-6 Astra" (September 3, 2026; GPT-6 Astra, GPT-5.6 Sol). Model names are those at the time of each announcement.
In the short band of the same table (4K–8K), GPT-5.5 scored 98.1% and GPT-5.4 97.3%. All are figures OpenAI published in its own announcements, not third-party measurements.
This isn't "the low-scoring models are bad." In the short band of the same table, GPT-5.5 and GPT-5.4 both score above 97%, and most real work — code review, long-form writing, meeting-note summaries, research synthesis — finishes well short of 1M. The problem is the "it has 1M, so just throw 1M at it" approach. Keep in mind, too, that these are figures OpenAI published in its own announcements, and they are only comparable within a table. Google also lists MRCR v2 (8-needle) results in its Gemini 3.1 Pro model card (February 2026) — 84.9% at 128K (average) versus 26.3% at 1M — but it splits the lengths differently, so it isn't among the bars above. Anthropic's announcement pages for Claude Opus 5.5 and its other current models don't list length-by-length scores of this kind. To know what the model you use today can do, check it against your own use case with the method below.
Measure the "effective range" for your own use case
Benchmark numbers come from document types and question styles that differ from yours. The most reliable approach is to build a small needle-in-a-haystack test with your own documents.
- Take the kind of document you actually use (code, meeting notes, contracts, and so on) and place 3–5 facts with clear answers at the beginning, middle, and end (facts already in the document work too)
- Ask it to "list all of them and say where each one appears," so it has to retrieve several facts at once. Asking for just one makes the model look better than it is (the single- vs multi-needle gap)
- Repeat the same question at different lengths — say 50K, 200K, 500K — and note the length where answers start to break down
- Include questions that combine facts, not just retrieve them ("compare A and B"). Questions that require synthesis tend to break down sooner
Just below the length where things start breaking is that model's "effective context" for that use case. Re-measure when you switch models. It takes tens of minutes, and it informs real decisions better than any spec-sheet number.
5. The Cost Trap — Models Whose Rates Jump on Long Inputs, and Models Whose Rates Don't
The last chapter said the effective range sits below the limit. On top of that comes a second trap: sending long inputs can make pricing jump. Vendors have designed this differently.
| Model (as of Sep 2026) | Standard rate (input / output, per 1M tokens) | Long inputs |
|---|---|---|
| Claude Opus 5.5 | $4 / $20 | Same rate across the full 1M |
| GPT-6 Sol | $2 / $10 | Inputs over 272K: the whole request at 2x input, 1.5x output |
| GPT-6 Astra | $10 / $50 | Same as above |
| Gemini 3.1 Pro (preview) | $2 / $12 | Over 200K: input $4, output $18 |
Let's run the numbers. Take a case where you send a 500K-token document and get one 50K-token response — the typical "summarize a large codebase or annual report in one go" scenario.
- Claude Opus 5.5 (flat): $4 × 0.5 + $20 × 0.05 = $3.00
- GPT-6 Sol (surcharge over 272K): $4 × 0.5 + $15 × 0.05 = $2.75 ($1.50 without the surcharge)
- Gemini 3.1 Pro (over-200K rate): $4 × 0.5 + $18 × 0.05 = $2.90 ($1.60 at the up-to-200K rate)
- GPT-6 Astra (surcharge over 272K): $20 × 0.5 + $75 × 0.05 = $13.75
There are two ways to read this. First, the moment you cross the threshold, the same model costs about 1.8x as much. GPT-6 Sol costs half as much as Claude Opus 5.5 on short inputs, but at 500K the gap nearly disappears — which model is "cheaper" flips with input length. Second, when a surcharge stacks on a high-priced flagship, the cost jumps to a different scale (GPT-6 Astra comes out at more than 4x Opus 5.5). Recalculate with the models you're comparing and your own typical input length before you choose.
With threshold pricing, split work that can be split so each piece stays below the threshold. Send the 500K as two 250K requests and GPT-6 Sol charges no surcharge — $1.50 total (though this doesn't work for tasks that need to see everything at once). It's the same structure I covered in "AI Token and Session Cost Saving."
6. Five Saving Tactics — Ranked by Real Impact for Solo Devs
"The container is 1M but the effective range ends well before that, and using it long gets expensive." We've covered that. So what can you actually do in the field? Here are five tactics I use day-to-day, ranked by what gives the biggest payoff.
Context Saving — Priority Order
/compact or start a new session.
Of the five, tactic ① "Cut the Session" gives the biggest visible gain. Just cutting the chat noticeably reduces hallucinations.
Tactic ④ is for API developers — UIs (claude.ai / ChatGPT) handle caching automatically.
My personal best practice: just doing ① and ② consistently shifts perceived accuracy noticeably. Even with Claude Code, instead of pushing one long session, hitting /compact or starting a fresh session at every topic change keeps final output quality stable.
Summary
Key points from this article:
- Context window = the maximum number of tokens an AI can handle in one exchange. The container size
- As of September 2026, the top models from Anthropic, OpenAI, and Google are all in the 1M-token class. Input limits barely differ; the differences are in output limits and long-input pricing
- The advertised limit and the effective range are different things. On OpenAI's published multi-needle benchmark MRCR (8 needles), the same model scores lower as the input grows, and scores near 1M ranged from 32.2% for Claude Opus 4.7 (in OpenAI's table) to 96.3% for GPT-6 Astra
- The most reliable way to find the effective range is to scatter facts through your own documents and test at different lengths. Re-measure when you switch models
- Long-input pricing comes in two shapes: the same rate up to the limit (Anthropic), or a surcharge past a threshold (272K for OpenAI, 200K for Google). Which is cheaper flips with input length
- Saving comes down to five moves: cut sessions, send excerpts, repeat instructions at the end, cache, and make addresses explicit — ① and ② matter most
The container got bigger, yet what we're really doing is still choosing what to pass in and what to leave out. Today's AI skill is no longer "the ability to cram everything in." It's the judgment to pass exactly what's needed, correctly — a skill that stays useful no matter how many model generations come and go. That's my conclusion now that 1M has become the norm.
FAQ
OpenAI has the tiktoken library, and Anthropic's API has a token counting feature. Rough guides: 1 Japanese character ≈ 1–1.5 tokens, 1 English word ≈ 1.3–1.8 tokens — but this shifts with the tokenizer generation (see the note in §5). Code varies a lot by type, too, so measuring before you send a long input is the safe move.
Context lives only inside the session — close the chat and it's gone. Memory (ChatGPT Memory / Claude Memory) is a separate cross-session retention mechanism. Memory contents end up injected into the context window, but from the user's perspective it's persistent vs. ephemeral.
RAG is the pattern of "dynamically fetch only the necessary information into context." Even with a 1M window, dumping everything makes it slow, heavy, and expensive, so retrieval-then-load (RAG) remains the mainstream approach. See What Is RAG for more.
Several factors stack up: a mismatch between the sequence lengths mostly seen in training and those used at inference, limits of positional encoding in the attention mechanism, and the steep rise in computation needed to combine multiple pieces of information. "Supported" and "accurate across the whole window" are separate problems. Where the drop begins depends on the model and the task, so the reliable way is to measure with your own documents using the method in §4.
Yes. MCP is a fetch-on-demand mechanism via tools, so you don't need to load everything into context up front. Switch the mental model from "paste the whole file" to "let it go read the file."
Decide by goal. ① If you want a summary of the whole thing, summarize each chapter first, then combine those summaries in a second pass. ② If you're looking for a specific answer, don't send the full text — pull out only the relevant parts with search (RAG). ③ Only when you need to compare across the entire document should you send it in one go, at a length that fits the effective range. When splitting, cut at headings and overlap a few paragraphs on either side so the thread doesn't break at the seams.