Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 2 days ago • 11
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection Paper • 2607.19942 • Published 3 days ago • 4
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published 10 days ago • 13
Trajectory-aware Cross-view Geo-localization with Sequential Observations Paper • 2607.15491 • Published 9 days ago • 6
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 4 days ago • 209
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 5 days ago • 53
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting Paper • 2607.18084 • Published 5 days ago • 5
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 6 days ago • 88
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 9 days ago • 67
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 9 days ago • 197
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 9 days ago • 53
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published 16 days ago • 74
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization Paper • 2607.04988 • Published 19 days ago • 27
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published 27 days ago • 22
Native Active Perception as Reasoning for Omni-Modal Understanding Paper • 2606.19341 • Published Jun 17 • 19
ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining Paper • 2606.17200 • Published Jun 15 • 55
MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold Paper • 2606.13376 • Published Jun 11 • 16