Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go? Paper • 2607.17986 • Published 30 days ago • 6
How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulation Paper • 2606.16821 • Published Jun 15 • 4
Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine Paper • 2510.21614 • Published Oct 24, 2025 • 22
PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors Paper • 2507.15550 • Published Jul 21, 2025 • 6