In 2023, a 32K-token context window felt "spacious." Today, a window in the 1-million-token (1M) class is simply assumed for top-tier models. As of September 2026, Anthropic, OpenAI, and Google all list an input limit of roughly one million tokens in the official specs of their top models. Individual model names and figures change with every release, so I leave those to the list of current models and cutoff dates and to each vendor's official pages. This article focuses on what survives a generation change: how to read the numbers, and how to work with them.

"One million tokens" translates to roughly 8–10 paperback books in English, or tens of thousands of lines of source code. We can now keep that much "in view" within a single session. But a document fitting in the container is not the same as the model reading all of it. The scores OpenAI publishes on its multi-needle long-context benchmark (MRCR) show that the same model scores lower as the input gets longer, and how far it drops varies widely by model and generation (covered in §1 and §4).

Let me get my take out front: the era of choosing a model purely on container size is over. What matters is the trio of "effective context × cost × how you feed it." This article walks through what context actually is, how to read a spec sheet, why bigger isn't enough, how to measure the effective range for your own use case, how long-input pricing jumps, and five saving tactics that solo developers and small teams can apply today — using official figures and published benchmark numbers.

CONTEXT WINDOW · 2023→2026

The Container Grew 250x in Three Years

— A timeline of how 1M went from luxury to baseline

2023
4K–200K
GPT-3.5 and early GPT-4 had 4K–32K — one research paper filled it. In November, Claude 2.1 shipped 200K.
2024
128K–2M
GPT-4 Turbo (128K) and Claude 3 (200K) became the norm. In June, Gemini 1.5 Pro opened 2M to developers.
2025
1M spreads
GPT-4.1 in April and Claude Sonnet 4 (beta) in August added 1M support.
2026
1M = standard
Top Claude, GPT, and Gemini models all sit in the 1M-token class (as of September).

But "supported" and "read to the end" are different things. On OpenAI's long-context benchmark MRCR (8 needles), GPT-5.5 goes from 98.1% at 4K–8K to 74.0% at 512K–1M.
How far scores fall differs by model and generation (evaluation table in OpenAI's "Introducing GPT-5.5," April 23, 2026; details in §1 and §4).

1. 1M Support Is Everywhere — But "Reads to the End" Is Another Matter

1M support spread fast over the past two years. In April 2025 OpenAI's GPT-4.1 shipped with about 1.05 million tokens, in August 2025 Anthropic's Claude Sonnet 4 reached 1 million tokens (in beta), and as of September 2026 the top models from Anthropic, OpenAI, and Google all list roughly one million tokens in their official specs. Back in 2023, 32K felt spacious — that's more than 30x in just three years. The container-size race looked finished.

But look at the long-context scores the vendors publish themselves, and the picture gets more complicated. The benchmark with a full set of results by length is OpenAI's MRCR v2 (8 needles). It hides eight requests of the same kind — say, "write a poem about tapirs" — inside a long conversation with an AI, then asks for one specific instance, such as "return the 2nd poem," exactly. Because the model has to tell near-identical items apart, down to their order, it is a multi-needle needle-in-a-haystack test. The score is how closely the returned text matches the correct one. Here are the results by length:

  • GPT-5.5: 98.1% at 4K–8K, 87.5% at 128K–256K, 74.0% at 512K–1M
  • GPT-5.4 (the previous generation in the same table): 97.3% → 79.3% → 36.6% across the same three bands
  • Claude Opus 4.7 (figures OpenAI listed in the same table): 59.2% at 128K–256K, 32.2% at 512K–1M
  • GPT-6 Astra (announced September 2026): 100.0% at 256K–512K, 96.3% at 512K–1M (GPT-5.6 Sol in the same table: 91.5% and 73.8%)

Sources (checked September 26, 2026): evaluation tables in OpenAI's "Introducing GPT-5.5" (April 23, 2026) and "GPT-6 Astra" (September 3, 2026). How the benchmark works: the OpenAI MRCR dataset card. Model names are those at the time of each announcement.

Two things stand out. First, the same model scores lower as the input gets longer. Second, how steeply it drops varies widely by model and generation — in the band closest to 1M, GPT-5.4 fell below 40%, while GPT-6 Astra, announced in September 2026, stayed above 90%. Rankings reshuffle with every generation, so these numbers go stale quickly. What lasts is the lesson: the advertised limit and the range where accuracy actually holds are two different numbers.

Don't misread this. It's not "Claude or GPT is bad." Use cases that truly need all 1M are rarer than you'd think. If a model reads stably up to 300K (roughly 2–3 books), almost every coding, research, and summarization task fits. The problem is choosing on the "1M supported" number alone — that's how you end up with the wrong decision criteria.

2. What Is Context? — Separate the Container from Its Contents

Quick terminology. Three words get mixed up in this space.

Three Terms

Token, Window, Context

① TOKEN — Unit of Text
The smallest unit AI processes text in. ~4 English characters per token (or ~0.75 of a word); CJK languages run roughly 1–1.5 tokens per character.
② WINDOW — Container Size
The maximum number of tokens a model can handle in a single exchange. Input plus output (including reasoning) combined. Via the API, input alone that exceeds it returns an error; chat apps and agents usually make room by summarizing or dropping older parts.
③ CONTEXT — The Contents
What's currently loaded in the window. Includes the system prompt, conversation history, attachments, tool outputs — all of it.

In short: "window = container size," "context = contents," "token = unit."
A big container with messy contents still gives you messy answers.

Also: don't confuse "context" with "memory." Context lives inside the session — close the chat and it's gone. Features like ChatGPT Memory or Claude Memory, on the other hand, are a separate cross-session retention mechanism. Memory contents do eventually get injected into the context window, but from the user's perspective it's persistent storage vs. ephemeral workspace.

Common misconception: "Bigger context window = smarter AI" is wrong. Window size is just the upper bound on what can be in view. Reasoning ability, knowledge depth, and instruction-following accuracy are measured separately. Every model release leads with "1M context!" as the headline, but that's only one facet of capability.

3. Read the Container Size from Three Numbers

When you look at a model's spec sheet, there are only three things to check around context. Get these down and you can compare any model with the same procedure.

① Input limit

The number the marketing leads with. But it tells you how much fits in, not how much actually gets read — as the next section covers, the effective figure is far below it.

② Output limit

Often overlooked, but it is an order of magnitude smaller than the input limit. Even among major models as of September 2026, you can send in 1 million tokens but get back only about 65K–128K. For "rewrite this whole long document" jobs, this is the limit you hit first.

③ Billing model

Flat across the whole window, or a threshold where the unit price jumps. This matters most in production and is usually the one thing missing from the spec sheet. §5 works the numbers out.

Below are the figures on each vendor's official pages as of September 29, 2026. The numbers move with each generation, so read this as a worked example of how the three axes above turn into real differences. For today's specific model names, see the list of current models and cutoff dates; for the numbers, check each vendor's official pages.

Family (examples as of Sep 2026)Input limitOutput limitLong-input pricing
Anthropic top tier (Claude Fable 5.1, Opus 5.5, Sonnet 5.5)1,000,000128,000Same rate all the way to the limit
Anthropic lightweight (Claude Haiku 4.5)200,00064,000—
OpenAI (GPT-6 Astra, Sol)1,050,000128,000Inputs over 272K: the whole request at 2x input, 1.5x output
Google (Gemini 3.1 Pro, preview)1,048,57665,536Over 200K: input $2→$4, output $12→$18

Sources (checked September 29, 2026): Anthropic Models overview and Pricing / OpenAI model pages for GPT-6 Astra and GPT-6 Sol / Google Gemini 3.1 Pro Preview and Gemini API pricing

Look at the table and the input limits are nearly identical across vendors; the differences are in output limits and the shape of the pricing. When the limits match, window size stops being a reason to choose. What does differ: Anthropic keeps the same rate up to the limit, while OpenAI and Google raise the rate past a certain length — so under the same "1M support" label, how freely you can send long inputs varies. That's not just a pricing detail; it reflects different approaches to long-context workloads. The cost chapter runs the numbers.

In practice, the choice goes like this. Decide by the document sizes you routinely work with. If they fit under ~200K, choose on accuracy stability in that band and whether you stay below a pricing threshold, not on window size. Only if you routinely handle giant documents over 300K do the width of the effective range and the long-input rate become deciding factors. Don't commit to one model — split by use case. That way of deciding holds no matter which generation is current.

Where to check the limits

Check the numbers on each vendor's official pages, not in roundup sites or article tables (including this one). Where to look is fairly consistent:

  • Anthropic: the comparison table on the models overview page has "Context window" and "Max output" rows. Long-input pricing is in the "Long context pricing" section of the pricing page
  • OpenAI: each model's page in the API docs lists "context window" and "max output tokens," plus a note on pricing for long prompts
  • Google: the Gemini API model page lists "Input token limit" and "Output token limit," and the pricing page has a "prompts > 200k tokens" tier

One more thing that's easy to miss: the same limit holds different amounts of text depending on the model. Limits are counted in tokens, so when the tokenizer (the scheme that splits text into tokens) changes, the token count of the same document changes too (see the note in §5). When you switch models, recount with your own documents to confirm you still have headroom.

4. Three Reasons "Bigger Is Better" Doesn't Hold

The spec table in the last chapter only shows the physical size of the container. So does the model actually use the container it advertises? Bottom line: don't assume it reads with the same accuracy all the way to the limit. There are three reasons.

Reason ①: Lost in the Middle

This is the phenomenon researchers from Stanford and elsewhere (Liu et al.) reported in 2023 in the paper "Lost in the Middle." On tasks such as finding an answer across multiple documents, models answered best when the answer sat at the beginning or end of the input and lost significant accuracy when it sat in the middle — even models built for long context showed the same pattern.

In practice it feels like this: "Paste a long PDF in full, ask 'What's the figure for X?', and it mixes up a number right around the middle." That's Lost in the Middle. How severe it is varies by model, but the safe assumption is that information placed mid-document is easier to miss — and to adjust how you feed it accordingly.

Reason ②: Context Rot

The longer a conversation runs, the more your early instructions fade. You asked for a formal tone at the start, and 20 turns later the model has drifted back to casual — that's Context Rot.

Two causes. ① Early instructions get treated as relatively old and light within the conversation history. ② The longer history spreads attention thin, making specific tokens harder to reference. In its developer documentation, Anthropic calls the drop in accuracy and recall as token counts grow "context rot," and writes that choosing what goes into context matters as much as how much fits. In September 2025 it laid out how to handle the problem as a deliberate skill in an engineering post titled "Effective context engineering for AI agents".

Reason ③: Advertised Context ≠ Effective Context

Take only the band closest to 1M (512K–1M) from the §1 figures and line them up. All come from OpenAI's announcement evaluation tables and use the same benchmark, OpenAI MRCR v2 (8 needles).

OpenAI MRCR v2 (8 needles) × 512K–1M

Near 1M, can the model return the one item asked for?

GPT-6 Astra Sep 2026 96.3%
GPT-5.5 Apr 2026 74.0%
GPT-5.6 Sol Sep 2026 73.8%
GPT-5.4 Apr 2026 36.6%
Claude Opus 4.7 Apr 2026 · OpenAI's table 32.2%

Sources: evaluation tables in OpenAI's "Introducing GPT-5.5" (April 23, 2026; GPT-5.5, GPT-5.4, Claude Opus 4.7) and "GPT-6 Astra" (September 3, 2026; GPT-6 Astra, GPT-5.6 Sol). Model names are those at the time of each announcement.
In the short band of the same table (4K–8K), GPT-5.5 scored 98.1% and GPT-5.4 97.3%. All are figures OpenAI published in its own announcements, not third-party measurements.

This isn't "the low-scoring models are bad." In the short band of the same table, GPT-5.5 and GPT-5.4 both score above 97%, and most real work — code review, long-form writing, meeting-note summaries, research synthesis — finishes well short of 1M. The problem is the "it has 1M, so just throw 1M at it" approach. Keep in mind, too, that these are figures OpenAI published in its own announcements, and they are only comparable within a table. Google also lists MRCR v2 (8-needle) results in its Gemini 3.1 Pro model card (February 2026) — 84.9% at 128K (average) versus 26.3% at 1M — but it splits the lengths differently, so it isn't among the bars above. Anthropic's announcement pages for Claude Opus 5.5 and its other current models don't list length-by-length scores of this kind. To know what the model you use today can do, check it against your own use case with the method below.

Measure the "effective range" for your own use case

Benchmark numbers come from document types and question styles that differ from yours. The most reliable approach is to build a small needle-in-a-haystack test with your own documents.

  1. Take the kind of document you actually use (code, meeting notes, contracts, and so on) and place 3–5 facts with clear answers at the beginning, middle, and end (facts already in the document work too)
  2. Ask it to "list all of them and say where each one appears," so it has to retrieve several facts at once. Asking for just one makes the model look better than it is (the single- vs multi-needle gap)
  3. Repeat the same question at different lengths — say 50K, 200K, 500K — and note the length where answers start to break down
  4. Include questions that combine facts, not just retrieve them ("compare A and B"). Questions that require synthesis tend to break down sooner

Just below the length where things start breaking is that model's "effective context" for that use case. Re-measure when you switch models. It takes tens of minutes, and it informs real decisions better than any spec-sheet number.

5. The Cost Trap — Models Whose Rates Jump on Long Inputs, and Models Whose Rates Don't

The last chapter said the effective range sits below the limit. On top of that comes a second trap: sending long inputs can make pricing jump. Vendors have designed this differently.

Model (as of Sep 2026)Standard rate (input / output, per 1M tokens)Long inputs
Claude Opus 5.5$4 / $20Same rate across the full 1M
GPT-6 Sol$2 / $10Inputs over 272K: the whole request at 2x input, 1.5x output
GPT-6 Astra$10 / $50Same as above
Gemini 3.1 Pro (preview)$2 / $12Over 200K: input $4, output $18

Let's run the numbers. Take a case where you send a 500K-token document and get one 50K-token response — the typical "summarize a large codebase or annual report in one go" scenario.

  • Claude Opus 5.5 (flat): $4 × 0.5 + $20 × 0.05 = $3.00
  • GPT-6 Sol (surcharge over 272K): $4 × 0.5 + $15 × 0.05 = $2.75 ($1.50 without the surcharge)
  • Gemini 3.1 Pro (over-200K rate): $4 × 0.5 + $18 × 0.05 = $2.90 ($1.60 at the up-to-200K rate)
  • GPT-6 Astra (surcharge over 272K): $20 × 0.5 + $75 × 0.05 = $13.75

There are two ways to read this. First, the moment you cross the threshold, the same model costs about 1.8x as much. GPT-6 Sol costs half as much as Claude Opus 5.5 on short inputs, but at 500K the gap nearly disappears — which model is "cheaper" flips with input length. Second, when a surcharge stacks on a high-priced flagship, the cost jumps to a different scale (GPT-6 Astra comes out at more than 4x Opus 5.5). Recalculate with the models you're comparing and your own typical input length before you choose.

With threshold pricing, split work that can be split so each piece stays below the threshold. Send the 500K as two 250K requests and GPT-6 Sol charges no surcharge — $1.50 total (though this doesn't work for tasks that need to see everything at once). It's the same structure I covered in "AI Token and Session Cost Saving."

Note: The same document can have a different token count on a different model. Limits and pricing are both counted in tokens, so a new tokenizer changes both the cost of a document and how much headroom you have. In its official models overview, Anthropic states that on the current tokenizer, used since Claude Opus 4.7, 1M tokens holds about 555K words, versus about 750K words on earlier models — roughly 1.35x as many tokens for the same English text. Even if the per-token rate stays the same, compare actual bills after switching.

6. Five Saving Tactics — Ranked by Real Impact for Solo Devs

"The container is 1M but the effective range ends well before that, and using it long gets expensive." We've covered that. So what can you actually do in the field? Here are five tactics I use day-to-day, ranked by what gives the biggest payoff.

Five Practical Tips

Context Saving — Priority Order

① Cut the Session
When the topic shifts, open a new chat. Just stopping the old context from carrying over eliminates Context Rot. In Claude Code, use /compact or start a new session.
② Send Excerpts, Not Full Texts
Pasting a 100-page PDF whole is the worst move. Use grep / search to extract relevant sections, compress to 3–5 pages, then send. The RAG mindset, applied solo.
③ Repeat Key Instructions at the End
Lost-in-the-Middle countermeasure. Restate the rule from the top in one line at the end: "Given the above, output in format X."
④ Prompt Caching
If you reuse the same system prompt or reference material, Anthropic/OpenAI prompt caching brings the input rate for the cached portion down to 10% of the base rate or less (official pricing as of September 2026). If you're calling the API, set this up first.
⑤ Make File Addresses Explicit
Specifying "file N, line X" boosts retrieval accuracy in long contexts. Think of it as handing the AI a table of contents with index entries.

Of the five, tactic ① "Cut the Session" gives the biggest visible gain. Just cutting the chat noticeably reduces hallucinations.
Tactic ④ is for API developers — UIs (claude.ai / ChatGPT) handle caching automatically.

My personal best practice: just doing ① and ② consistently shifts perceived accuracy noticeably. Even with Claude Code, instead of pushing one long session, hitting /compact or starting a fresh session at every topic change keeps final output quality stable.

Summary

Key points from this article:

  • Context window = the maximum number of tokens an AI can handle in one exchange. The container size
  • As of September 2026, the top models from Anthropic, OpenAI, and Google are all in the 1M-token class. Input limits barely differ; the differences are in output limits and long-input pricing
  • The advertised limit and the effective range are different things. On OpenAI's published multi-needle benchmark MRCR (8 needles), the same model scores lower as the input grows, and scores near 1M ranged from 32.2% for Claude Opus 4.7 (in OpenAI's table) to 96.3% for GPT-6 Astra
  • The most reliable way to find the effective range is to scatter facts through your own documents and test at different lengths. Re-measure when you switch models
  • Long-input pricing comes in two shapes: the same rate up to the limit (Anthropic), or a surcharge past a threshold (272K for OpenAI, 200K for Google). Which is cheaper flips with input length
  • Saving comes down to five moves: cut sessions, send excerpts, repeat instructions at the end, cache, and make addresses explicit — ① and ② matter most

The container got bigger, yet what we're really doing is still choosing what to pass in and what to leave out. Today's AI skill is no longer "the ability to cram everything in." It's the judgment to pass exactly what's needed, correctly — a skill that stays useful no matter how many model generations come and go. That's my conclusion now that 1M has become the norm.

FAQ

Q1. How can I measure token counts in advance?

OpenAI has the tiktoken library, and Anthropic's API has a token counting feature. Rough guides: 1 Japanese character ≈ 1–1.5 tokens, 1 English word ≈ 1.3–1.8 tokens — but this shifts with the tokenizer generation (see the note in §5). Code varies a lot by type, too, so measuring before you send a long input is the safe move.

Q2. How is "memory" different from context?

Context lives only inside the session — close the chat and it's gone. Memory (ChatGPT Memory / Claude Memory) is a separate cross-session retention mechanism. Memory contents end up injected into the context window, but from the user's perspective it's persistent vs. ephemeral.

Q3. How does RAG relate to the context window?

RAG is the pattern of "dynamically fetch only the necessary information into context." Even with a 1M window, dumping everything makes it slow, heavy, and expensive, so retrieval-then-load (RAG) remains the mainstream approach. See What Is RAG for more.

Q4. Why does accuracy drop well before a 1M window is full?

Several factors stack up: a mismatch between the sequence lengths mostly seen in training and those used at inference, limits of positional encoding in the attention mechanism, and the steep rise in computation needed to combine multiple pieces of information. "Supported" and "accurate across the whole window" are separate problems. Where the drop begins depends on the model and the task, so the reliable way is to measure with your own documents using the method in §4.

Q5. Do MCP servers save context?

Yes. MCP is a fetch-on-demand mechanism via tools, so you don't need to load everything into context up front. Switch the mental model from "paste the whole file" to "let it go read the file."

Q6. How should I split or summarize a document too long for the effective range?

Decide by goal. ① If you want a summary of the whole thing, summarize each chapter first, then combine those summaries in a second pass. ② If you're looking for a specific answer, don't send the full text — pull out only the relevant parts with search (RAG). ③ Only when you need to compare across the entire document should you send it in one go, at a length that fits the effective range. When splitting, cut at headings and overlap a few paragraphs on either side so the thread doesn't break at the seams.