laya-multilingual-rknn

Multilingual Laya decision scorer for Rockchip RK3588. One fixed FP16 RKNN graph, sequence length 512, up to 64 option markers. Built with RKNN-Toolkit2 2.3.2. Weights are not quantized (do_quantization=False).

The graph scores one question per run. Logits come back for every marker slot; slots past the real options are filled with -10000. Divide by temperature.json before the softmax if you want probabilities. All three temperatures in this export are 1.0, so the softmax is unchanged.

Files

File What it is
laya-multilingual-fp16-seq512.rknn FP16 RKNN graph, RK3588, toolkit 2.3.2
tokenizer.json mmBERT tokenizer from the multilingual checkpoint
temperature.json Per-question-type temperature, all 1.0
scripts/ ONNX export, logit-mask rewrite, RKNN build

librknnrt.so is not in this repo. Load librknnrt 2.3.2 from the RKNN-Toolkit2 v2.3.2 release, and point the process at that file (RKNN_LIB or an explicit dlopen). Runtime 2.4.2 loads this graph and then returns a constant vector on an RK3588 with NPU driver 0.9.8. Do not replace the system librknnrt.so if another process needs a different runtime.

Inputs

The graph inputs are the same five tensors Laya feeds the scorer, padded to the compiled shape:

Name Shape dtype
input_ids 1 × 512 int64
attention_mask 1 × 512 int64
marker_pos 1 × 64 int64
marker_mask 1 × 64 int64
qtype 1 int64

qtype is 0 choice, 1 score, 2 noul. Output logits is 1 × 64 float. Token layout matches laya.Agent for the multilingual checkpoint (scripts/export_onnx.py uses that code path).

On device, set the NPU core mask to all cores (RKNN_NPU_CORE_ALL, 0xFFFF). The default mask picks one idle core. Attention, norm, and GELU stay on core 0 in toolkit 2.3.2, so three cores do not triple the speed. A full 512-length run is about 500 ms on one core and about 510–527 ms with all three. A real 65-token prompt is still padded to 512.

These ops stay on CPU: the int64 inputs, the Equal that builds the attention mask, the token-embedding Gather (256000 × 768, about 10 ms), and GatherElements for the markers (about 8 ms). Toolkit 2.3.2 does not place that embedding table on the NPU.

Check

Against PyTorch fp16 on 88 live Laya questions: argmax matched on 87 / 88, mean logit cosine 0.999. The miss is a satisfaction question whose top two options differ by 0.002; fp16 swaps them. That is one comparison, not a new benchmark of the original model.

Provenance

Piece Source
Checkpoint convaiinnovations/laya, subfolder multilingual (mmBERT-base encoder)
Backbone license jhu-clsp/mmBERT-base, MIT
Export laya==0.3.20, transformers==5.17.0, eager attention, ONNX opset 18
Graph edit Final logit mask rewritten from Where/Equal to Mul/Add
RKNN RKNN-Toolkit2 2.3.2, target_platform=rk3588, float_dtype=float16, optimization_level=3, disable_rules=["seperate_conv_split"], do_quantization=False

Toolkit 2.4.2 was tried on the same ONNX. The simulator matched PyTorch. The device, with runtime 2.4.2, did not. The file in this repo is the 2.3.2 build.

Convert it again

On an x86 machine with Python 3.10. The RK3588 board only runs the finished .rknn.

python3.10 -m venv .venv
.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu
.venv/bin/pip install 'laya==0.3.20' 'transformers==5.17.0' onnx
# RKNN-Toolkit2 2.3.2 wheel from the v2.3.2 GitHub release (cp310)
.venv/bin/pip install rknn_toolkit2-2.3.2-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

.venv/bin/python scripts/export_onnx.py --out work/laya_multilingual_seq512.onnx
.venv/bin/python scripts/patch_logit_mask.py work/laya_multilingual_seq512.onnx work/laya_mask_only.onnx
.venv/bin/python scripts/build_rknn.py work/laya_mask_only.onnx laya-multilingual-fp16-seq512.rknn

export_onnx.py also rewrites temperature.json from the loaded checkpoint.

License

Apache-2.0, the license of the Laya checkpoint. mmBERT-base is MIT; its notice is in NOTICE. The Rockchip runtime library is under Rockchip's own terms and is not part of this Apache-2.0 distribution.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ShiWarai/laya-multilingual-rknn

Finetuned
(152)
this model