Your Agents Don't Need More Memory. They Need a Better Budget.
DeepMind's Dream-RSI froze every model weight and rewrote the search loop instead — and quietly showed that feeding agents 'lessons learned' makes them worse.
Your Agents Don't Need More Memory. They Need a Better Budget.
Buried in the middle of a paper Google DeepMind dropped on arXiv last Monday is a result that should end about a dozen product roadmaps. The researchers tried the thing everyone in this industry has been selling since January — distill an agent's history into high-level insights, inject them back as prompt guidance — and it consistently made performance worse. Same budget. Same task. Guided agent loses.
At Kuaray, we spend a lot of time talking clients out of building "agent memory layers." Now we have a citation.
What Dream-RSI Actually Does
Dream-RSI (Zheng et al., Sept 14) is a recursive self-improvement framework, but not the kind the term usually implies. Nothing gets retrained. Per the authors: "Only the exploration-policy code changes; the underlying models, evaluator, and execution interfaces remain fixed."
The loop is almost embarrassingly pragmatic:
- A coding agent explores a task, building a discovery tree — each node a workspace snapshot plus a score.
- Completed trees become a replay simulator. Candidate policies get re-run against recorded history, deterministically. No new generations, no new spend.
- The winner gets promoted. Next round, repeat.
The policy being optimized is just orchestration code: which branch to pursue, how many workers to fan out, when to stop. It never invents new solutions during replay — it only re-decides how much to spend and where.
The Number That Matters Isn't the Benchmark
Read the tables honestly and Dream-RSI's quality wins are thin. It beats its own fixed-exploration baseline on the Lasso average — but loses on five of six individual datasets. On the autocorrelation-inequality task it loses outright to the baseline it was supposed to improve.
The cost column is a different story.
| Comparison | Result |
|---|---|
| vs. SimpleTES (Lasso) | ~162× fewer agent calls |
| vs. fixed exploration (Lasso) | 1.7× fewer calls, better average |
| KernelBench VGG16 | Comparable output, 2.43× fewer generations |
| KernelBench ConvDiv | 2.09× higher performance, same budget |
That's the headline for anyone with an agentic line item on next year's P&L. The frontier here is not smarter. It's cheaper at parity — and the learned policy got there by doing something no human would have signed off on: cutting evaluated attempts from 110 down to 50 mid-run, then spending hard again once progress plateaued.
Your platform team wrote a fixed max_parallel = 10 and hasn't touched it since March. That constant is now a line item.
Three Things to Do About It
- Instrument your agent runs as trees, not logs. Parent state, artifact, score, per node. If you can't replay a run, you can't optimize what it cost. Most teams have unstructured traces and no score at all.
- Stop building the insight extractor. §5.1 is unambiguous. Strong semantic priors over-constrain the search. Spend that sprint on evaluation signal instead.
- Treat orchestration constants as tunable, not architectural. Fan-out, depth, stopping rule. Those three numbers move your agentic bill more than your model choice does.
The uncomfortable bit for engineering leadership: this entire result was achievable with an LLM rewriting a Python file and a decent scoring function. No training run. No new model. Just someone bothering to measure what their agents were spending and letting the loop adjust.
Schedule a Technical Architecture Review with our Strategists — and bring your agent invoices.