Back to articles

Context Language Models: When Agents Edit Their Own Context

October 2, 2026
Context Language Models: When Agents Edit Their Own Context

Kuaray Deep Dive 02: Meta and UW's Context Language Models let agents rewrite their own working memory. The results, what to use today, and the cache edit tax.

Share:LinkedInX

Your context is a file now

Kuaray Deep Dive #02 — Context Language Models: the model stops just appending text and starts editing its own working memory. What the paper shows, what you can use today, and the cache math nobody is doing.

The problem: agents drown in their own history

Every long-running agent hits the same wall. At each step, the tool observation, the reasoning and the answer get appended to the context. Over a task that runs for hours, the history grows until it overflows the window, gets expensive, or dilutes the model's attention with stale noise.

Today's default fix is a fixed harness: once the context crosses a threshold, an external process summarizes everything and starts over. That's what Codex and most frameworks do. It works, but it's a human-written rule applied blindly: it summarizes at the wrong moment, throws away the detail that mattered, and keeps the log that no longer does.

On September 29, 2026, researchers from Meta Superintelligence Labs and the University of Washington (with MIT and Trillium Labs) proposed something else: let the model decide [1].

Part 1 — What a Context Language Model is

A traditional language model treats context as append-only. Each output is concatenated to what was already there: c(t+1) = c(t) ⊕ output.

A Context Language Model (CLM) replaces that with a free transformation: c(t+1) = f(c(t)). The context becomes an editable file, mirrored and synced with the model server, which the agent can rewrite with ordinary bash commands [1]:

  • append, insert, delete and replace segments in the middle of the context;
  • define reusable compaction functions and call them again later;
  • coordinate multiple context files, one per subagent, in multi-agent systems.

The most striking part is the behavior that emerged on its own. One agent kept a scoreboard inside its context, with 163 edits over the task, holding the context between 6K and 8K tokens. Others invented an internal "notes" role, filtered their history in loops, and wrote helper functions to compact themselves [1][3].

Append-only vs CLM

Part 2 — Two ways to get a CLM

The paper shows you can get there without training anything, or with training.

1. By prompting (in-context learning). Natural-language instructions teach the model to manage its context. A textual evolution loop improves the instruction document: the agent runs, a proposer model suggests better versions of the document, they're scored on a dev split and tested on a held-out split. It works with Claude 4.6 Sonnet and GPT-5.6-Sol through the API, and the proposer can be a stronger model (Fable 5.1) or the agent itself (self-evolution) [1][2].

2. By reinforcement learning. For smaller models, the authors use stepwise GRPO with a clever twist: a success-gated efficiency advantage. The model only earns credit for being economical among trajectories that solved the task. That blocks the obvious shortcut of deleting important context just to spend less [1].

Part 3 — The numbers

The results use prefix-reuse FLOPs, a metric that counts not only generation but the cost of re-processing the context after every mid-context edit. It's the honest metric for this kind of system [1].

TaskCLMBest baselineCompute
BrowseComp-Plus (deep research)59.4%53.4% (Codex-style summarization)−21.5% FLOPs
TBLite (terminal coding)73.7%67.0%91% of FLOPs
TerminalBench 2.1tiesbest baseline70% of FLOPs
EdgeBench-10 (single repo, 12 h)44.642.3−59% (179 vs 437 PFLOPs)
Software World (6 agents, multiple repos, 24 h)+65% speedupsummary-based swarmsame compute
Qwen3.5-9B with RL (BrowseComp-Plus)28.8% → 42.5%42.1% (summarization)~39% fewer FLOPs

On open-ended math, with Claude 4.6 Sonnet and a 32K context, the CLM was competitive with specialized evolutionary workflows (OpenEvolve) and came out ahead on circle packing (2.636 vs 2.541) [1].

On the server side, a new technique called Suffix Cache Reuse (SCR) reuses the KV cache after a mid-context edit, not just the prefix. The result: 35% less server compute than standard SGLang at the same accuracy [1].

Results

Part 4 — The math nobody is doing: the edit tax

Here's the detail that matters if you run on a hosted API, and it ties back to our Deep Dive 01 on LLM gateways and routing on cost.

Prompt caching works by prefix. On Anthropic's API, cache reads cost 0.1× the input price (0.05× on Opus 5.5), and cache writes cost 1.25× (5-minute TTL) [5]. An edit in the middle of the context invalidates the cache from that point on: everything after the edit gets written again instead of read [4][5].

SCR fixes this on the server, but hosted APIs don't yet support cache reuse after mid-context edits [3]. So every edit has a price. You can work out when it pays for itself:

Break-even turns ≈ S × (write − read) ÷ (D × read) − 1 S = tokens after the edit point · D = tokens removed

For a 32K context on Claude Sonnet 5.5 ($2 input · $2.50 cache write · $0.20 cache read per million):

Where the edit happensRemovesTokens afterBreaks even inExtra cost of the edit
Near the end8K2K~2 turns$0.003
Right after the pinned header (compaction)20K8K~4 turns$0.014
In the middle8K12K~16 turns$0.026
In the middle, on Opus 5.58K12K~35 turns$0.056

Edit tax

Three practical lessons:

  1. Where you edit matters more than how much you delete. The same 8K-token removal pays off in 2 turns near the end and in 16 in the middle.
  2. Models with cheaper cache reads charge more per edit. On Opus 5.5, reads are so cheap that the same edit takes twice as many turns to pay off.
  3. An expired cache changes everything. If the agent spends more than 5 minutes between calls (a slow tool, a human in the loop approval), the cache is already gone and the edit is practically free.

This math is cost only. It ignores the paper's main win: a shorter, better-organized context also gets more right. But it explains why harnesses that work always follow the same rule: pin the header, edit near the end, batch your edits.

Part 5 — What you can use today

You don't have to wait for a new model:

  • Anthropic API, context editing (server-side). The clear_tool_uses_20250919 strategy clears old tool results past a trigger (default: 100K tokens), keeps the last N, and has a clear_at_least parameter that exists precisely to make sure a clear removes enough tokens to justify invalidating the cache. There's also clear_thinking_20251015 and server-side compaction, compact_20260112 [4].
  • Memory tool. Combined with context editing, Claude gets a warning before clearing and can save what matters to memory files it can look up later [4].
  • The paper's harness. Meta published the code at facebookresearch/context-language-models [1].
  • Frameworks adopting it. There's already a proposal for an opt-in CLM mode in Hermes Agent, with the system prompt, task and memory pinned (non-editable), nudges at 25/50/75% of the budget, and batched edits to preserve the cache [6].

A good starting point in production: context in three zones.

ZoneWho editsExample
Pinnednobodysystem prompt, rules, task definition
Working memorythe modelscoreboard, notes, plan, summaries
Tailappend-onlylatest turns and tool results

Part 6 — The risks

  • Self-injection. If the model rewrites its own context, a malicious instruction from a web page or document can get "promoted" to a permanent note and survive many turns. The paper cites documented cases of models inserting unauthorized instructions into summaries and proposes no defenses, leaving that as future work [1][3].
  • Deleting what mattered. Without supervision, the model can discard exactly the detail it will miss 50 turns later. The paper itself admits outcome-only rewards are a weak signal for editing decisions [1].
  • Auditability. A context that changes mid-stream is harder to debug and explain. Preventing the loss of critical information and auditing changes are still open questions [7].
  • Benchmark not yet public. ContextBench, where the gains of up to 35.9 points show up, hasn't been released, which blocks independent verification [3].
  • No memory across tasks. A CLM organizes context within a task. When the task ends, it discards everything. Persistent memory is still a separate system [3].

Mitigations we recommend: a pinned zone the model can't touch; version every edit as a logged diff; filter instruction-like lines the model writes into working memory [6]; and evals that measure not just the final answer but what got deleted along the way. Silent context loss is one of the failure modes we catalogued in why AI agents fail in production.

Conclusion

For years we treated context like a log: it only grows, and something outside cleans it up. A CLM treats it as a working file the model organizes itself. The paper's numbers are strong: up to 59% less compute on a 12-hour task, and 65% more output from a 24-hour swarm.

If you run agents on a hosted API, the immediate lesson is an engineering one: every edit has a cache price, and where you edit decides whether it pays off. Pin the header, edit near the end, batch the changes, and track cache_read_input_tokens vs cache_creation_input_tokens on every call.

At Kuaray Tech, we build agents that manage their own memory, with clear rules about what they can and can't delete. That's core to our AI and agent engineering work, and if you're starting from zero, our MVP approach puts the context architecture in before the first long-running task.

Talk to Kuaray about your agent's context strategy — we audit how your agents grow and compact their context, and model what each edit really costs on your API bill.

References

  1. Shao, R. et al. — Context Language Models. arXiv:2609.37725v1 (Sept 29, 2026). Meta Superintelligence Labs, University of Washington, MIT, Trillium Labs. https://arxiv.org/html/2609.37725v1
  2. Mpost — Meta Presents Context Language Models (Oct 1, 2026). https://mpost.io/meta-presents-context-language-models-ai-agents-that-edit-their-own-memory-outperform-fixed-harnesses-at-lower-compute-cost/
  3. DEV Community (G. Dadhich) — Context Language Models: What the UW and Meta Paper Changes for Agent Builders, and What It Leaves to Memory (Oct 1, 2026). https://dev.to/gaurav_dadhich/context-language-models-what-the-uw-and-meta-paper-changes-for-agent-builders-and-what-it-leaves-40a7
  4. Anthropic — Context editing (accessed Oct 1, 2026). https://platform.claude.com/docs/en/build-with-claude/context-editing
  5. Anthropic — Prompt caching (cache write and read pricing, accessed Oct 1, 2026). https://platform.claude.com/docs/en/build-with-claude/prompt-caching
  6. NousResearch / hermes-agent — Issue #129584: CLM-style model-editable context as opt-in compression mode. https://github.com/NousResearch/hermes-agent/issues/129584
  7. CCTest — Context Language Models: Self-Managing AI Context. https://cctest.ai/en/articles/context-language-models-let-models-manage-their-own-context
  8. Aran Komatsuzaki (X) — paper announcement. https://x.com/arankomatsuzaki/status/2105181276714242518

Paper published on September 29, 2026. Numbers are the authors', not yet independently replicated. The "edit tax" math is ours, using Anthropic list prices at the time of writing.

Share:LinkedInX