view article Article Welcome RL Environments to the hub +6 burtenshaw, AdithyaSK, sergiopaniego, xeophon, ryanmarten, merve, lhoestq, julien-c • 10 days ago • 25
view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 16 days ago • 101
view article Article Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community +3 pcuenq, lysandre, victor, julien-c, Jundot • 16 days ago • 88
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper • 2606.19348 • Published Apr 26 • 47
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models Paper • 2508.06471 • Published Aug 8, 2025 • 217
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 28 days ago • 174
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published Sep 1 • 567
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • Aug 21 • 68
view article Article We changed one line and the benchmark score moved 0.21 AUROC FINAL-Bench • Aug 22 • 15
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • Aug 18 • 118
SWE-bench Collection SWE-bench (Lite, Verified, Multimodal, Multilingual) all in one place! • 5 items • Updated Dec 14, 2025 • 17