vedant33 's Collections

Supersede: Memory-Update Gap in LLM Agents

Open RL environment where the reward is temporal fact-currency. GRPO-trained Qwen2.5-3B LoRA lifts held-out supersession 9.0 -> 16.7 percent.