Tag: benchmark
All the articles with the tag "benchmark".
-
The three regimes of agentic KV eviction
On real Claude Code traces under memory pressure, LRU does 60 percent more recompute than a strong offline oracle. Three 2026 papers propose policies to close that gap. Measured against the oracle on shared traces, most do not beat LRU; the one signal that reliably helps, lifecycle retirement, does so only in a middle band of pressure and only by a variable amount. Here is the map, with the error bars.