Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Abstract
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field around that axis of weights versus skills. Its central analytical contribution is a deep-dive that arranges code-as-policy methods by their degree of self-improvement, from zero-shot program synthesis, through closed-loop self-repair and persistent skill memory, to the sparsely populated cell in which execution feedback, skill memory, and evolutionary search combine into one open-ended loop; only a few very recent systems (for example ASPIRE, ENPIRE, and RoboClaw) occupy that cell. We map the complementary "skills" pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and show that the word "skill" is used in at least five distinct senses, of which only the code sense self-improves without gradient updates. We then connect the taxonomy to the emerging skill economy: commercial robot-skill marketplaces now distribute one-tap skills across robots but ship only static playback, which surfaces open problems of adaptation, cross-embodiment portability, provenance, safety verification, composition, and standardisation. This is a deliberately focused survey. Rather than cataloguing the field exhaustively, it examines 77 representative systems across six technique families through one taxonomy and a set of contrast tables, and it supplies operational definitions of the self-improvement mechanisms together with a statement of what each family cannot do.
Community
This survey reframes modern robot learning around what actually shipsβfrozen policy weights versus executable skillsβand introduces a five-rung taxonomy of code-as-policy systems based on increasingly powerful combinations of execution feedback, persistent memory, and program search, culminating in autonomous self-improving robot skill loops.
β‘οΈ πππ² ππ’π π‘π₯π’π π‘ππ¬ π¨π ππ‘π πππ’π π‘ππ¬-π―π¬-ππ€π’π₯π₯π¬ π π«ππ¦ππ°π¨π«π€:
π§ πΎππππππ ππ. πΊπππππ π»πππππππ: Introduces a unified taxonomy spanning 77 core systems across six robot-learning familiesβcode-as-policy, end-to-end VLA, reward synthesis, skill libraries, sim-to-real/transfer, and benchmarksβplus 225 landscape works. Its key architectural distinction is whether competence is encoded in frozen neural weights (e.g., VLA backbone + action head) or represented as inspectable, executable programs/skills that can be edited and recombined after deployment. The taxonomy on page 3 makes this decomposition explicit.
π ππππ-πΉπππ πΊπππ-π°ππππππππππ π³ππ π ππ (π + π΄ + πΊ): The paper's main analytical novelty is decomposing code-as-policy agents by three operational mechanismsβFeedback (F) from execution, Memory (M) persisted across tasks, and Search (S) over multiple candidate programsβand arranging systems from zero-shot synthesis β closed-loop repair β skill-library accumulation β evolutionary search β full F+M+S self-improvement. Crucially, it distinguishes sequential debugging from genuine search and frozen model parameters from runtime memory, making βself-improvementβ technically testable rather than a loose label.
π§ ππππ πΊπππ-π°ππππππππ πΉππππ π³πππ + πΊππππ π¬ππππππ: Identifies a sparsely populated frontierβrepresented by ASPIRE, ENPIRE, and RoboClawβwhere an agent executes skills, obtains grounded traces/feedback, stores validated skills in persistent memory, and searches/mutates candidate programs, feeding accumulated competence into future tasks. The architecture is summarized in the page-11 diagram as Actor Agent β Execution Engine (F) β Skill Memory (M) β Evolutionary Search (S) with skills recursively returned to future tasks. The survey argues this is the missing adaptation layer between today's static robot-skill marketplaces and genuinely deployable, continually improving robot ecosystems.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- ASPIRE: Agentic /Skills Discovery for Robotics (2026)
- A Few Words Go a Long Way: Language Guided Robot Policy Synthesis (2026)
- RHO: Your Coding Agent is Secretly a Roboticist (2026)
- Sequential Planning via Anchored Robotic Keypoints (2026)
- SkillPlug: Unsupervised Skill Mining for Few-Shot Adaptation in Robotic Manipulation (2026)
- LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories (2026)
- ETA: A New Agentic Paradigm for Embodied Tasks (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper