Instructions to use ShiWarai/laya-multilingual-rknn with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Laya
How to use ShiWarai/laya-multilingual-rknn with Laya:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
laya-multilingual-rknn
Multilingual Laya decision scorer for Rockchip RK3588. One fixed FP16 RKNN graph, sequence length 512, up to 64 option markers. Built with RKNN-Toolkit2 2.3.2. Weights are not quantized (do_quantization=False).
The graph scores one question per run. Logits come back for every marker slot; slots past the real options are filled with -10000. Divide by temperature.json before the softmax if you want probabilities. All three temperatures in this export are 1.0, so the softmax is unchanged.
Files
| File | What it is |
|---|---|
laya-multilingual-fp16-seq512.rknn |
FP16 RKNN graph, RK3588, toolkit 2.3.2 |
tokenizer.json |
mmBERT tokenizer from the multilingual checkpoint |
temperature.json |
Per-question-type temperature, all 1.0 |
scripts/ |
ONNX export, logit-mask rewrite, RKNN build |
librknnrt.so is not in this repo. Load librknnrt 2.3.2 from the RKNN-Toolkit2 v2.3.2 release, and point the process at that file (RKNN_LIB or an explicit dlopen). Runtime 2.4.2 loads this graph and then returns a constant vector on an RK3588 with NPU driver 0.9.8. Do not replace the system librknnrt.so if another process needs a different runtime.
Inputs
The graph inputs are the same five tensors Laya feeds the scorer, padded to the compiled shape:
| Name | Shape | dtype |
|---|---|---|
input_ids |
1 × 512 |
int64 |
attention_mask |
1 × 512 |
int64 |
marker_pos |
1 × 64 |
int64 |
marker_mask |
1 × 64 |
int64 |
qtype |
1 |
int64 |
qtype is 0 choice, 1 score, 2 noul. Output logits is 1 × 64 float. Token layout matches laya.Agent for the multilingual checkpoint (scripts/export_onnx.py uses that code path).
On device, set the NPU core mask to all cores (RKNN_NPU_CORE_ALL, 0xFFFF). The default mask picks one idle core. Attention, norm, and GELU stay on core 0 in toolkit 2.3.2, so three cores do not triple the speed. A full 512-length run is about 500 ms on one core and about 510–527 ms with all three. A real 65-token prompt is still padded to 512.
These ops stay on CPU: the int64 inputs, the Equal that builds the attention mask, the token-embedding Gather (256000 × 768, about 10 ms), and GatherElements for the markers (about 8 ms). Toolkit 2.3.2 does not place that embedding table on the NPU.
Check
Against PyTorch fp16 on 88 live Laya questions: argmax matched on 87 / 88, mean logit cosine 0.999. The miss is a satisfaction question whose top two options differ by 0.002; fp16 swaps them. That is one comparison, not a new benchmark of the original model.
Provenance
| Piece | Source |
|---|---|
| Checkpoint | convaiinnovations/laya, subfolder multilingual (mmBERT-base encoder) |
| Backbone license | jhu-clsp/mmBERT-base, MIT |
| Export | laya==0.3.20, transformers==5.17.0, eager attention, ONNX opset 18 |
| Graph edit | Final logit mask rewritten from Where/Equal to Mul/Add |
| RKNN | RKNN-Toolkit2 2.3.2, target_platform=rk3588, float_dtype=float16, optimization_level=3, disable_rules=["seperate_conv_split"], do_quantization=False |
Toolkit 2.4.2 was tried on the same ONNX. The simulator matched PyTorch. The device, with runtime 2.4.2, did not. The file in this repo is the 2.3.2 build.
Convert it again
On an x86 machine with Python 3.10. The RK3588 board only runs the finished .rknn.
python3.10 -m venv .venv
.venv/bin/pip install torch --index-url https://download.pytorch.org/whl/cpu
.venv/bin/pip install 'laya==0.3.20' 'transformers==5.17.0' onnx
# RKNN-Toolkit2 2.3.2 wheel from the v2.3.2 GitHub release (cp310)
.venv/bin/pip install rknn_toolkit2-2.3.2-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
.venv/bin/python scripts/export_onnx.py --out work/laya_multilingual_seq512.onnx
.venv/bin/python scripts/patch_logit_mask.py work/laya_multilingual_seq512.onnx work/laya_mask_only.onnx
.venv/bin/python scripts/build_rknn.py work/laya_mask_only.onnx laya-multilingual-fp16-seq512.rknn
export_onnx.py also rewrites temperature.json from the loaded checkpoint.
License
Apache-2.0, the license of the Laya checkpoint. mmBERT-base is MIT; its notice is in NOTICE. The Rockchip runtime library is under Rockchip's own terms and is not part of this Apache-2.0 distribution.
- Downloads last month
- -
Model tree for ShiWarai/laya-multilingual-rknn
Base model
convaiinnovations/laya