Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
Paper • 2606.27472 • Published
Open RL environment where the reward is temporal fact-currency. GRPO-trained Qwen2.5-3B LoRA lifts held-out supersession 9.0 -> 16.7 percent.
Note Paper
Note GRPO LoRA adapter (9.0 -> 16.7%)
Note RL episodes (train + held-out test)
Supersession demo, GRPO-trained Qwen2.5-3B
Note Interactive supersession demo