-
Code as Agent Harness
Paper • 2605.18747 • Published • 225 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 195 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 171 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
Collections
Discover the best community collections!
Collections including paper arxiv:2605.00658
-
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 86 -
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Paper • 2604.28185 • Published • 92 -
Representation Fréchet Loss for Visual Generation
Paper • 2604.28190 • Published • 32 -
Co-Evolving Policy Distillation
Paper • 2604.27083 • Published • 68
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 31 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 45 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 24
-
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Paper • 2509.24832 • Published -
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 86 -
Map2World: Segment Map Conditioned Text to 3D World Generation
Paper • 2605.00781 • Published • 25 -
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
Paper • 2604.24026 • Published • 22
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 155 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 148 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 157 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 150
-
Code as Agent Harness
Paper • 2605.18747 • Published • 225 -
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Paper • 2605.12500 • Published • 195 -
From Context to Skills: Can Language Models Learn from Context Skillfully?
Paper • 2604.27660 • Published • 171 -
PhysBrain 1.0 Technical Report
Paper • 2605.15298 • Published • 145
-
SemShareKV: Efficient KVCache Sharing for Semantically Similar Prompts via Token-Level LSH Matching
Paper • 2509.24832 • Published -
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 86 -
Map2World: Segment Map Conditioned Text to 3D World Generation
Paper • 2605.00781 • Published • 25 -
From Skill Text to Skill Structure: The Scheduling-Structural-Logical Representation for Agent Skills
Paper • 2604.24026 • Published • 22
-
UniVidX: A Unified Multimodal Framework for Versatile Video Generation via Diffusion Priors
Paper • 2605.00658 • Published • 86 -
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
Paper • 2604.28185 • Published • 92 -
Representation Fréchet Loss for Visual Generation
Paper • 2604.28190 • Published • 32 -
Co-Evolving Policy Distillation
Paper • 2604.27083 • Published • 68
-
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
Paper • 2603.25746 • Published • 155 -
TAPS: Task Aware Proposal Distributions for Speculative Sampling
Paper • 2603.27027 • Published • 148 -
Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
Paper • 2603.25716 • Published • 157 -
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Paper • 2603.27538 • Published • 150
-
EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters
Paper • 2402.04252 • Published • 31 -
Vision Superalignment: Weak-to-Strong Generalization for Vision Foundation Models
Paper • 2402.03749 • Published • 15 -
ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Paper • 2402.04615 • Published • 45 -
EfficientViT-SAM: Accelerated Segment Anything Model Without Performance Loss
Paper • 2402.05008 • Published • 24