rlmath agentic GRPO checkpoints

Intermediate checkpoints from agentic GRPO runs on rlmath, a collection of math construction and optimization environments with deterministic verifiers and no LLM judge, packaged as Harbor tasks. The agent is Terminus-2. It works in a sandboxed terminal: it writes and runs code, then submits a construction that a programmatic grader scores.

Each folder is a standalone Hugging Face model directory with safetensors weights, config, tokenizer and chat template. It loads with from_pretrained(<repo>, subfolder="<folder>"). Optimizer states are not included.

Common setup

  • Trainer: verl (fully-async policy, partial rollout) with the Alibaba agentic recipe (remote agent loop and LLM proxy). vLLM rollouts, FSDP training.
  • Algorithm: GRPO, lr 1e-6 (constant), clip 0.2 / 0.28, no KL loss, token-mean loss aggregation, temperature 1.0, one optimizer step per training step.
  • Reward: construct tasks give 1/0. Optimize tasks give 0 if the submission is invalid, otherwise 0.1 + 0.9 · clipped progress toward the target.
  • Validation: one sample per task at temperature 0.6, top-p 0.95. Two validation sets were used:
    • v1 (contaminated): 139 tasks, one per family. They were sampled from the training pool, and 93 of the 139 are also in the 1,145-task training set used from rlmath_async_v2 onward. These numbers partly measure training tasks.
    • v2 (held out): 121 tasks across 48 families with no overlap with the 1,024-task training set tasks_train_v2. All runs from the "v2 recipe" section below use it.

Checkpoints

v1 recipe (validation set v1, contaminated)

folder base run settings validation
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_25 Qwen3-4B-Thinking-2507 16 prompts × 8 rollouts, 8 turns, 8k tokens/call, 32k total, 3,401 tasks step 25: 0.235
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_50 〃 〃 —
qwen3-4b-thinking/rlmath_async_v2/step_25 v1 step 25 1,145 non-trivial tasks, overlong filtering, entropy bonus 0.001 step 0: 0.235 → step 25: 0.134
qwen3-4b-thinking/rlmath_4b_kt_v1/step_36, step_48 Qwen3-4B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.054 (attempt 1, step 0) / 0.227 / 0.294 / 0.299 / 0.306 at steps 0 / 12 / 24 / 36 / 48
qwen3-30b-a3b-thinking/rlmath_30b_v1/step_6 Qwen3-30B-A3B-Thinking-2507 16 × 8, 8 turns, 12k/call, 32k total, 1,145 tasks step 0: 0.341
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v1/step_36, step_48 Qwen3-30B-A3B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.331 / 0.372 / 0.447 / 0.521 / 0.503 at steps 0 / 12 / 24 / 36 / 48

v2 recipe (validation set v2, held out)

Shared settings: 1,024 training tasks, 16 prompts × 32 rollouts per step, up to 12 turns, 32k tokens per model call, 128k tokens per trajectory, 192 planned steps (3 passes). Trajectories cut off by a length limit are dropped from the loss, and generation stops once a trajectory's budget is spent. Token-level truncated importance sampling (cap 2.0) corrects for the vLLM-vs-trainer mismatch, and each turn after the first is prompted with exactly the token sequence the trainer trains on. Weights are bf16 with torchao AdamW (bf16 stochastic rounding) under FSDP1.

folder base run-specific settings held-out validation (v2)
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v2/step_36, step_48 Qwen3-30B-A3B-Thinking-2507 4 vLLM + 4 trainer GPUs, Ulysses SP 2, staleness 0.5, entropy bonus 0.001 0.427 / 0.479 / 0.548 / 0.536 / 0.539 at steps 0 / 12 / 24 / 36 / 48
qwen3.5-4b/rlmath_q35_4b_kt_v1/step_24, step_36, step_156, step_168 Qwen3.5-4B 5 vLLM + 2 trainer GPUs, staleness 1.0, entropy bonus 0; run stopped at step ~171 of 192 (sandbox outage), step_156 is the best checkpoint 0.399 / 0.323 / 0.483 / 0.529 / 0.544 / 0.441 / 0.540 / 0.561 / 0.518 / 0.532 / 0.508 / 0.483 / 0.564 / 0.599 / 0.547 at steps 0 / 12 / 24 / … / 156 / 168 (every 12)
qwen3.5-4b/rlmath_q35_4b_jupiter_v1/step_6, step_12, step_14, step_18, step_20 Qwen3.5-4B one 4×GH200 node (2 vLLM + 2 trainer GPUs), staleness 1.0, entropy bonus 0.001 0.343 / 0.511 at steps 0 / 12 (validated every 12 steps only)

Qwen3.5 notes. Qwen3.5 is a hybrid Gated-DeltaNet / attention model and is natively multimodal. Training was text-only: the vision tower is included in these folders but was never updated. Micro-batches hold one sequence each, because transformers' Gated-DeltaNet does not respect packed-sequence boundaries.

Entropy blowup in rlmath_q35_4b_jupiter_v1. With the 0.001 entropy bonus, token entropy rose monotonically from about 0.66 (step 6) to 0.91 (step 12), 1.16 (step 14), 1.49 (step 18) and 1.82 (step 20), while responses grew to 30–51k tokens. A Qwen3.5-9B run with the same bonus collapsed after the same pattern (entropy 3.6, train reward halved by step 32). Step 12 is the last checkpoint before the blowup; steps 14–20 are included for analysis and are not recommended for use. The KT run of the same model uses no entropy bonus and stayed stable.

researchmath2 (held-out researchmath2 validation)

rlmath_q35_4b_rm2_jupiter_v1 trains on a different task set: researchmath2 (research-level construct / optimize / short-form problems derived from arXiv papers). It uses 2,587 training tasks and a 128-task validation set, split by source paper, so no validation paper is seen in training. The sandbox image adds research math tools (SageMath, GAP, Singular, Macaulay2, PARI/GP, nauty, SAT/SMT solvers); on JUPITER it is an aarch64 Apptainer port of rlmath-base:0.3-tools.

Settings: Qwen3.5-4B, 16 prompts × 32 rollouts, up to 12 turns, 32k tokens per call, 128k per trajectory, entropy bonus 0, truncated trajectories dropped from the loss, and trials that failed for infrastructure reasons left out of their GRPO group. The rest is as in the v2 recipe. It ran on 4 JUPITER GH200 nodes, with 8 vLLM and 8 trainer GPUs, and 323 steps (2 passes) are planned. Training is still running.

folder held-out validation (researchmath2, 128 tasks)
qwen3.5-4b/rlmath_q35_4b_rm2_jupiter_v1/step_12 … step_120 (every 12), step_128 see below

Validation used two protocols, and only numbers within one protocol are comparable:

  • Steps 0–24: temperature 0.6, and an episode ends after its first truncated model call. About 55–60% of episodes ended that way after one over-long first reply, so these numbers mostly measure that. Scores: 0.097 / 0.156 / 0.158 at steps 0 / 12 / 24.
  • Step 28 onward: temperature 1.0, top-p 0.95, and no early stop in validation (the training-time stop still applies). Scores: 0.324 / 0.342 / 0.460 / 0.431 / 0.458 / 0.539 / 0.439 / 0.501 / 0.510 / 0.596 at steps 28 / 36 / 48 / 60 / 72 / 84 / 96 / 108 / 120 / 132. The step-132 checkpoint was not saved; step_128 is the latest uploaded.

Known issue in the earlier 4B runs

rlmath_q4bthink_async_v1 and rlmath_async_v2 were trained with a train/rollout context mismatch. Qwen3-Thinking's chat template drops the reasoning of earlier assistant turns, so turns after the first were sampled without that reasoning in context but trained with it (vLLM-vs-trainer KL ≈ 0.045). Both runs degraded after about 25 steps. All later runs (rlmath_4b_kt_v1, rlmath_30b_*, rlmath_q35_*) prompt every turn with exactly the token sequence the trainer sees, which brings the KL down to about 1e-3.

The 30B runs use bf16 weights with torchao's AdamW (bf16 stochastic rounding) under FSDP1.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for amphora/rlmath-agentic-grpo-checkpoints

Finetuned
(44)
this model