SKILL-KD: Contrastive Skill Distillation for LLM Agents Paper • 2607.28048 • Published 3 days ago • 8
Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking Paper • 2607.00482 • Published 3 days ago • 6
ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads Paper • 2608.02703 • Published 4 days ago • 5
Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 6 days ago • 11
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published 1 day ago • 20
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation Paper • 2608.03632 • Published 3 days ago • 18
RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction Paper • 2608.01247 • Published 5 days ago • 10
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models Paper • 2608.03457 • Published 3 days ago • 27
MemSFT: Mitigating Alignment Tax with an External Parametric Memory Paper • 2607.25614 • Published 10 days ago • 22
PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning Paper • 2608.01837 • Published 4 days ago • 37
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning Paper • 2608.04007 • Published 3 days ago • 17
Zero-Mem: Zero-Token Memory Operations for LLM Agents Paper • 2607.29377 • Published 7 days ago • 9
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published 4 days ago • 29
GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning Paper • 2608.02585 • Published 4 days ago • 23
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 12 days ago • 99
MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis Paper • 2607.27146 • Published 9 days ago • 27
Flux-OPD: On-Policy Distillation with Evolving Contexts Paper • 2607.28022 • Published 8 days ago • 43