laya-finetuned-rlcd

A fine-tuned version of the Apache-2.0 laya multilingual decision model, trained with RLCD (Resampling for Local Calibration Distribution) on the MIT-licensed jev-playground-rlcd-v0 decision corpus.

Read this section first. The headline numbers below are measured against the soft-teacher's argmax, not independent human ground truth, on a 4,000-row training subset. They are strong evidence of improvement but should be read as "teacher-agreement accuracy + calibration," not as a human-verified decision-quality benchmark. See Limitations for the full caveat.

What it does

laya is a non-autoregressive System-1 decision model. Give it a state (an email, ticket, or JSON blob) and typed questions, and it returns calibrated probabilities in a single forward pass (~33 ms) across 100+ languages. It never generates text, so there's nothing to parse and nothing to hallucinate.

This checkpoint is fine-tuned for better accuracy and, more importantly, better calibration β€” a model whose reported probabilities are honest.

Results

Independent held-out set: 75,311 items (27,496 choice / 27,498 score / 20,317 noul), constructed as the full jev gold corpus minus the 4,000 training rows, so it is genuinely out-of-distribution. Both models run through the identical preprocess/forward path, so the comparison is apples-to-apples.

Type n BASE FINETUNED Ξ” BASE raw ECE FT raw ECE BASE Brier FT Brier BASE temp FT temp
choice 27,496 0.451 0.984 +0.533 0.302 0.137 16,751 1,378 6.70 1.07
score 27,498 0.173 0.988 +0.815 0.505 0.133 22,690 1,353 10.0 1.10
noul 20,317 0.647 0.885 +0.238 0.220 0.037 7,381 1,448 5.58 1.20
overall 75,311 0.403 0.959 +0.556

Two takeaways:

  1. Accuracy roughly doubled (0.403 β†’ 0.959). The score task is the most dramatic: base scored below random (0.173 on a 4-way task); the fine-tuned model scores 0.988.
  2. Calibration collapsed toward ideal. RLCD fits one temperature per question type. Base needs 5.6–10.0 (its logits are badly overconfident). The fine-tuned model needs ~1.07–1.20 β€” essentially no correction. A temperature near 1.0 is the RLCD target: the model's own probabilities are trustworthy.

Brier (a proper scoring rule; lower is better) improves ~13–16Γ— across all types.

Loading this checkpoint

This is a laya-format checkpoint, not a standard HF transformers model. It is loaded via the laya Python package, not via AutoModel.from_pretrained(). Upload the full laya checkpoint structure below.

pip install laya
from laya.agent import Agent

# Load the fine-tuned checkpoint directly from a local path, or clone the repo first:
a = Agent(model_id_or_path="path/to/laya-finetuned-rlcd")

q = {"type": "choice",
     "instructions": "Classify\n\"Please cancel my subscription.",
     "criteria": ["keep", "cancellation"]}
out = a.predict(state="x", questions={"q": q})
# out["answers"]["q"]["probabilities"] -> {"keep": ..., "cancellation": ...}

questions must be a dict {qid: {type, instructions, criteria}}; the output container is answers (not results).

Files in this repo

File Purpose
model.safetensors Fine-tuned weights (643 MB)
rl_agent_config.json Model config, incl. fitted temperatures [1.071, 1.034, 1.045]
encoder/ mmBERT/ModernBert encoder
tokenizer/ Tokenizer

Training recipe

Base convaiinnovations/laya (Apache-2.0), multilingual
Data jev-playground-rlcd-v0 (MIT); 4,000-row decision subset
Objective RLCD β€” policy-gradient loss over noisy logit projections (GRPO-style, 4 samples/item, exploration noise annealed 0.4β†’0.1), rewarded by proper scoring rules (spherical 0.75, ranked-probability 1.0); plus full-weight soft cross-entropy
Epochs 4
Optimizer AdamW, cosine schedule, WD 0.01
LRs encoder 2.5e-5, head 1e-4
Batch micro_batch 8, group_size 4, max_tokens_per_batch 4096
Lengths max_len 1024, head_max_len 256
Device CPU

Reproduce with the laya finetune script (/opt/laya-git/research/scripts/finetune_single_device.py or its checkpointed --resume variant).

Limitations

  1. Accuracy is teacher-agreement, not ground truth. The target is the soft teacher's argmax β€” agreement with training labels, not an independent human check. "Perfect on the teacher" is strong evidence but not a guarantee of correctness on reality.
  2. Subset, not full corpus. This checkpoint was trained on 4,000 jev rows. The full corpus (~31,500 rows) is available; a full-corpus run is the natural next step.
  3. Slight residual overconfidence. Fine-tuned temperatures are ~1.07–1.20, not exactly 1.0 β€” the model nudges slightly overconfident, highest on the noul task (1.20). Not a problem, but stated honestly.

Attribution

Downloads last month
48
Safetensors
Model size
0.3B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Snugasabug/laya-finetuned-rlcd

Finetuned
(161)
this model