Skip to content

A million-token window is not memory: three levers you can measure on your agent bill

A 1M-token window is not memory. In Chroma's tests accuracy drops in cliffs long before it fills; Anthropic's compaction summarizes at 150K and other vendors reprice above 200K–272K. In our view, the levers are a measured context budget, effort per call, and a filter that never calls the model.

retrieval · guardrails

A million-token context window is not memory. It is a ceiling, and in published tests quality falls off long before you reach it. If you pay the bill for an agent, in our view three levers matter more than the price per token: how much context you let in, how hard the model is told to think, and how much traffic never reaches the model at all. You can measure each one on your own system.

That is the answer. Below is the proof, then how we would set each lever.

What the measurements say

Context degrades in cliffs. Chroma Research's Context Rot report (July 2025) ran its tests on a sample of 18 frontier models. As a secondary summary of the report puts it, the work found large accuracy gaps between a focused 300-token prompt and the same question placed inside 113,000 tokens. The drop did not arrive as a gentle ramp. It came as cliffs, and it got worse when the distracting text looked like the answer. The counter-intuitive part: coherent, well-structured input hurt attention more than shuffled input, on all 18 models.

A separate comparison of long-context prompting against retrieval, published on usewire.io, looked at cost per query and at accuracy by where the relevant passage sits. It puts retrieval at roughly 1,250× cheaper per query, and reports long context losing more than 30% accuracy when the relevant passage sits in the middle of the window. We did not run either study.

The vendors design around it. Per the same summary, Anthropic's compaction defaults to summarizing at 150K tokens on a 1M-token model. OpenAI reprices requests above 272K input tokens, Google above 200K. In our reading, that is the plainest signal available: the companies that sell the window do not run it full.

Effort moves the bill. A single engineering breakdown on YouTube ran a real task, plan to implementation to verification, and forced a flagship model down to low reasoning effort. Against the high-effort run:

                 high effort     low effort (flagship)
time             14 min 24 s     7 min 55 s
input tokens     222,000         91,000
output tokens    32,000          11,000
tool calls       92              33

Two caveats. It is one breakdown of one task, not an industry benchmark, and we treat it that way. And the write-up we have is not fully clear on whether the high-effort baseline ran the same model or the previous generation. Until that is settled, we read the table as a direction, not a size.

A filter at both edges. This routing pattern is not ours, and we could not trace it to one named, dated report with a stated sample. Read it as a pattern reported from production agent deployments that hold up, not as a measurement. Those deployments put a hard deterministic filter on both ends of the queue. Alerts below severity 5 close automatically with no model call. Alerts above severity 12 bypass the agent and go to a human. The model gets only the ambiguous middle, 5 to 12. The same reports argue that an unbounded queue in front of an agent fails twice at once, a bloated bill on trivia and real risk on high-stakes calls, and that the guardrail belongs in routing logic outside the model, not in the prompt. We think that reading is right, which is why the filter gets its own lever.

Lever 1: a context budget you measured

Our view: treat the window size as a hardware limit, not a design target. The number that matters is the budget where accuracy on your own corpus still holds. That is an experiment, not a spec sheet.

# context budget: a sketch
for size in [2k, 8k, 32k, 64k, 128k]:
    for q in eval_questions:              # questions with known answers, from your corpus
        ctx = distractors(size, similar_to=q.answer)
        ctx.insert(position="middle", text=relevant_passage(q))
        record(size, correct = model(ctx, q) == q.answer)

safe_budget = largest size before the first cliff
# above safe_budget: retrieve, do not paste

Two details carry the weight. Make the distractors resemble the answer, because that is where Chroma saw degradation get worse. And put the passage in the middle of the window, where the usewire.io comparison reports the 30%+ loss. Our expectation is that a test with random filler and the answer at the top will look better than your production traffic does.

Long context does replace retrieval in one bounded case, in our view: a single document that fits inside the budget you measured.

Lever 2: effort per call type

The breakdown points one way, but with its baseline unclear we do not lean on its exact numbers. Our bet is on the direction: on routine calls, high effort mostly buys extra passes rather than better answers. That is an opinion to test, not a finding. So set effort per step, not per agent, and measure the change on the same model.

effort_by_step = {
    "classify_ticket":      "low",
    "extract_fields":       "low",
    "plan_change":          "high",
    "verify_against_tests": "medium",
}
# same model, one step at a time
# log per step: wall time, input tokens, output tokens, tool calls
# keep a lower setting only if your eval still passes

Log the four numbers the breakdown tracked. Tool calls deserve their own column. Every call pulls more text back into context, which feeds straight into lever 1.

Lever 3: the filter the model never sees

The cheapest model call is the one you never make. That is our opinion, and the edge pattern under it is the unattributed one described above. In code:

def route(alert):
    if alert.severity < 5:
        return auto_close(alert)      # no model call
    if alert.severity > 12:
        return page_human(alert)      # the agent never sees it
    return agent.triage(alert)        # the ambiguous middle only

Three lines do two jobs. The low edge takes volume off the bill. The high edge takes the dangerous decisions away from the model. Neither job can live in the prompt. "Escalate anything critical" is a request the model can misread. An if is not. The thresholds 5 and 12 are the example values in that pattern; yours should come from your own history.

How the levers compound

They are not independent. The filter cuts the number of calls. The measured budget caps tokens per call. Lower effort, where your eval allows it, cuts passes and tool calls inside each call, and fewer tool calls mean less text flowing back into the window. Our reading is that most "the agent got expensive" stories are one of these three left at its default.

What would change our mind: a measurement on a real corpus where accuracy holds flat to the full window, with answer-like distractors and the passage mid-window. We have not seen one.

FAQ

Should we drop retrieval now that windows reach 1M tokens? Not on this evidence. In Chroma's tests on 18 frontier models, accuracy fell in cliffs far below 1M. The usewire.io comparison puts retrieval at roughly 1,250× cheaper per query. And the vendors draw their own lines well short of 1M: Anthropic's compaction summarizes at 150K, while OpenAI and Google reprice above 272K and 200K.

Does low reasoning effort hurt quality? The breakdown we cite reports time, tokens and tool calls, not quality, and its baseline may not be the same model. Run your own eval before and after on the same model, and lower effort one step at a time.

Where should our filter thresholds come from? Severity 5 and 12 are the example values in the production pattern described above, which we could not trace to a named report. They are not a standard. In our view, set yours from your own history: what a human always ignored goes below the low edge, what a human must always own goes above the high one.

Isn't well-structured input safer than a raw dump? Not automatically. In Chroma's tests, coherent, well-structured input hurt attention more than shuffled input on all 18 models. Our takeaway: for the model, less text beats tidier text.

A related case

Holt, our language-learning product, is built on the same split as lever 3: the engine that decides what a learner sees next never calls a model, and models only write the content. The case is at iloblique.com/holt.