view article Article MiniMax Goes Sparse: Decoding M3's Attention from a Single Diagram AtlasCloud-AI • May 29 • 11
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 20 days ago • 76
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published 11 days ago • 137
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 49
WARP: Weight-Space Analysis for Recovering Training Data Portfolios Paper • 2607.01686 • Published 27 days ago • 10
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 7 days ago • 31
AutoIndex: Learning Representation Programs for Retrieval Paper • 2607.18603 • Published 8 days ago • 10
Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction Paper • 2605.11354 • Published May 12 • 2