NaFlexCLAP Collection Experiments in OpenCLIP (main branch) with `timm` NaFlexViT audio encoder + modern text encoder for variable-time , variable-length text CLAP models • 4 items • Updated 4 days ago • 3
LTX-2.5 Collection LTX-2.5 base models, quantized models and accompanying LoRAs and IC-LoRAs • 4 items • Updated 3 days ago • 39
You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences Paper • 2606.15956 • Published Jun 14 • 13
view article Article Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident +2 hlarcher, XciD, raphael-gl, chris-rannou • 19 days ago • 468
EO-Robotics Collection EmbodiedOneVision is a unified framework for multimodal embodied reasoning and robot control, featuring interleaved vision-text-action pretraining. • 7 items • Updated Mar 2 • 9
Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model Paper • 2607.03509 • Published Jul 3 • 14
From Foundation to Application: Improving VLA Models in Practice Paper • 2607.06403 • Published Jul 7 • 20
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory Paper • 2605.17478 • Published May 17 • 1
VISReg: Variance-Invariance-Sketching Regularization for JEPA training Paper • 2606.02572 • Published Jun 1 • 4
Duality Models: An Embarrassingly Simple One-step Generation Paradigm Paper • 2602.17682 • Published Feb 4 • 1
MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines Paper • 2603.06679 • Published Mar 30 • 7
view article Article Arcee Becomes the First Major American AI Lab to Replace AWS S3 with Hugging Face Private Storage, in a Multi-Million Dollar Commercial Partnership clem • Jun 9 • 35
view article Article How an Agent Built a 3D Paris Gallery by Chaining Two Hugging Face Spaces mishig • Jun 9 • 23