Instructions to use DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16 with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16") config = load_config("DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MiMo-V2.6-Distill-Qwen-9B — MLX BF16
Full BF16 MLX conversion. No quantization.
No quantization overrides.
Converter-reported size: BF16 weights. mlx-vlm did not print a bits/weight figure for this build.
What this is
Format conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B at commit f2773fb482ac3dd047a4af4003b86e56b7225d0d into MLX safetensors. Produced with mlx-vlm 0.7.2 and mlx 0.32.2 (CUDA 13 wheel) on spark-d500 (NVIDIA GB10). The four source shards matched the Hub LFS SHA-256 values before conversion.
This is not a new training run. Upstream describes the checkpoint as a supervised fine-tune of Qwen/Qwen3.5-9B on MiMo-generated agent data (code, general agent tasks, visual coding, cybersecurity). Upstream benchmark numbers were not re-run here.
The pinned upstream card does not declare a license. This repository redistributes converted weights of that public checkpoint.
Architecture
Qwen3_5ForConditionalGeneration/model_type: qwen3_5- Text: 32 layers, hidden 4096, 16 query heads / 4 KV heads, full attention every 4th layer, config context 262144
- Vision: Qwen3.5 vision tower, depth 27, hidden 1152, patch 16
- Chat template is the upstream MiMo v2.6 template shipped in
chat_template.jinja
Quantization
Mixed recipes are mlx-vlm's built-in predicates, not a sensitivity search. Group size 64, affine mode. The predicate skips multimodal modules, so the vision tower stays BF16.
The 22 high-bit overrides, read from config.json on mixed-4-6 and the same pattern on the other mixed builds, are:
embed_tokensandlm_headdown_projin layers 0, 1, 2, 3, 6, 9, 12, 15, 18, 21, 24, 27, 28, 29, 30, 31v_projonly in full-attention layers 3, 15, 27, 31
The other 228 quantized modules use the lower width. Those 22 / 228 figures are quantization-override entries, not safetensor tensor counts. Mixed builds store 1260 tensors because quantized weights are split into weight, scales, and biases. BF16 stores 760 tensors.
On mixed builds, top-level quantization.bits is 4. That is mlx-vlm's default field. The per-module bits entries are what was applied.
Smoke
CUDA smoke on spark-d500, device gpu:0, mlx-vlm 0.7.2. One greedy generation per build, max_tokens=64, temperature=0.0. Prompts went through apply_chat_template(..., enable_thinking=False). The formatted text prompt ended in <think></think> and had no image token. The image prompt contained <|vision_start|><|image_pad|><|vision_end|>.
Text prompt: What is 15% of 240? Answer with the number only.
Image-input smoke, mixed-4-6 only: a 64x64 solid red PNG, What color is the square in the image? Answer with one word.
| Build | Kind | Pass | Output |
|---|---|---|---|
bf16 |
text | yes | 36 |
mixed-3-5 |
text | yes | <value>36</value> |
mixed-3-6 |
text | yes | thinking block, then 36 |
mixed-3-8 |
text | yes | thinking block, then 36 |
mixed-4-6 |
text | yes | 36 |
mixed-4-6 |
image-input | yes | Red |
mixed-4-8 |
text | yes | 36 |
mixed-3-6 and mixed-3-8 emitted a <thinking> block even though the prompt closed thinking, then the number. mixed-3-5 wrapped the number in <value> tags. That is a quality difference on this one prompt, not a load failure. No perplexity or benchmark was run. Decode rates and peak memory are not reported: the BF16 call was cold, later outputs were 3–64 tokens, and peak memory stayed at the BF16 process high-water mark.
The pass/output log is smoke-results.json in this repo.
Usage
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
model, processor = load("DJLougen/MiMo-V2.6-Distill-Qwen-9B-MLX-bf16")
prompt = apply_chat_template(
processor,
model.config,
"What is 15% of 240? Answer with the number only.",
num_images=0,
enable_thinking=False,
)
print(generate(model, processor, prompt, max_tokens=64, temperature=0.0).text)
For an image, pass num_images=1 and the image path to generate. Image-input smoke was only run for mixed-4-6.
Provenance
| Item | Value |
|---|---|
| Source | XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Source revision | f2773fb482ac3dd047a4af4003b86e56b7225d0d |
| Converter | mlx-vlm 0.7.2, mlx 0.32.2 CUDA 13 |
| Machine | spark-d500, NVIDIA GB10 |
| Source shard check | SHA-256 matched Hub LFS hashes for all 4 shards |
| Smoke | smoke-results.json from the conversion host, 2026-09-21 |
- Downloads last month
- 133
Quantized