MiMo-V2.6-Distill-Qwen-9B — MLX BF16

Full BF16 MLX conversion. No quantization.

No quantization overrides.

Converter-reported size: BF16 weights. mlx-vlm did not print a bits/weight figure for this build.

What this is

Format conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B at commit f2773fb482ac3dd047a4af4003b86e56b7225d0d into MLX safetensors. Produced with mlx-vlm 0.7.2 and mlx 0.32.2 (CUDA 13 wheel) on spark-d500 (NVIDIA GB10). The four source shards matched the Hub LFS SHA-256 values before conversion.

This is not a new training run. Upstream describes the checkpoint as a supervised fine-tune of Qwen/Qwen3.5-9B on MiMo-generated agent data (code, general agent tasks, visual coding, cybersecurity). Upstream benchmark numbers were not re-run here.

The pinned upstream card does not declare a license. This repository redistributes converted weights of that public checkpoint.

Architecture

  • Qwen3_5ForConditionalGeneration / model_type: qwen3_5
  • Text: 32 layers, hidden 4096, 16 query heads / 4 KV heads, full attention every 4th layer, config context 262144
  • Vision: Qwen3.5 vision tower, depth 27, hidden 1152, patch 16
  • Chat template is the upstream MiMo v2.6 template shipped in chat_template.jinja

Quantization

Mixed recipes are mlx-vlm's built-in predicates, not a sensitivity search. Group size 64, affine mode. The predicate skips multimodal modules, so the vision tower stays BF16.

The 22 high-bit overrides, read from config.json on mixed-4-6 and the same pattern on the other mixed builds, are:

  • embed_tokens and lm_head
  • down_proj in layers 0, 1, 2, 3, 6, 9, 12, 15, 18, 21, 24, 27, 28, 29, 30, 31
  • v_proj only in full-attention layers 3, 15, 27, 31

The other 228 quantized modules use the lower width. Those 22 / 228 figures are quantization-override entries, not safetensor tensor counts. Mixed builds store 1260 tensors because quantized weights are split into weight, scales, and biases. BF16 stores 760 tensors.

On mixed builds, top-level quantization.bits is 4. That is mlx-vlm's default field. The per-module bits entries are what was applied.

Smoke

CUDA smoke on spark-d500, device gpu:0, mlx-vlm 0.7.2. One greedy generation per build, max_tokens=64, temperature=0.0. Prompts went through apply_chat_template(..., enable_thinking=False). The formatted text prompt ended in <think></think> and had no image token. The image prompt contained <|vision_start|><|image_pad|><|vision_end|>.

Text prompt: What is 15% of 240? Answer with the number only. Image-input smoke, mixed-4-6 only: a 64x64 solid red PNG, What color is the square in the image? Answer with one word.

Build Kind Pass Output
bf16 text yes 36
mixed-3-5 text yes <value>36</value>
mixed-3-6 text yes thinking block, then 36
mixed-3-8 text yes thinking block, then 36
mixed-4-6 text yes 36
mixed-4-6 image-input yes Red
mixed-4-8 text yes 36

mixed-3-6 and mixed-3-8 emitted a <thinking> block even though the prompt closed thinking, then the number. mixed-3-5 wrapped the number in <value> tags. That is a quality difference on this one prompt, not a load failure. No perplexity or benchmark was run. Decode rates and peak memory are not reported: the BF16 call was cold, later outputs were 3–64 tokens, and peak memory stayed at the BF16 process high-water mark.

The pass/output log is smoke-results.json in this repo.

Usage

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16")
prompt = apply_chat_template(
    processor,
    model.config,
    "What is 15% of 240? Answer with the number only.",
    num_images=0,
    enable_thinking=False,
)
print(generate(model, processor, prompt, max_tokens=64, temperature=0.0).text)

For an image, pass num_images=1 and the image path to generate. Image-input smoke was only run for mixed-4-6.

Provenance

Item Value
Source XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Source revision f2773fb482ac3dd047a4af4003b86e56b7225d0d
Converter mlx-vlm 0.7.2, mlx 0.32.2 CUDA 13
Machine spark-d500, NVIDIA GB10
Source shard check SHA-256 matched Hub LFS hashes for all 4 shards
Smoke smoke-results.json from the conversion host, 2026-09-21
Downloads last month
133
Safetensors
Model size
9B params
Tensor type
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support