Muse-Glimmer-30B-DFlash2

Blog | GitHub

This repository contains the DFlash 2 draft model for meta-models/Muse-Glimmer-30B. It is not a standalone language model: it runs inside a speculative decoding server and drafts tokens for the target model to verify. It is finetuned from meta-models/Muse-Glimmer-30B-assistant, the official DFlash drafter Meta ships with the model. This repository is a mirror of incoai/Muse-Glimmer-30B-DFlash2.

DFlash 2 is a block-diffusion drafter for speculative decoding. It predicts a whole block of tokens in a single pass and keeps the top candidates at every position. A lightweight selector then traces one coherent path through them. Two-tap dynamic convolutions in the backbone keep the draft from decaying toward the end of the block. Decoding is lossless: greedy output matches the target model exactly, and sampling preserves its distribution.

DFlash 2: parallel block drafting with a candidate path selector

Quick Start

Serve with SGLang:

pip install "sglang[all] @ git+https://github.com/sgl-project/sglang.git#subdirectory=python"

python -m sglang.launch_server \
  --model-path meta-models/Muse-Glimmer-30B \
  --speculative-algorithm DFLASH \
  --speculative-draft-model-path incoai/Muse-Glimmer-30B-DFlash2 \
  --speculative-num-draft-tokens 16

Or with vLLM:

pip install -U "vllm @ git+https://github.com/vllm-project/vllm.git@refs/pull/52816/head"

vllm serve meta-models/Muse-Glimmer-30B \
  --speculative-config '{
    "method": "dflash",
    "model": "incoai/Muse-Glimmer-30B-DFlash2",
    "num_speculative_tokens": 15
  }'

See the blog post for other engines and more details.

Evaluation

  • Runtime: SGLang on one NVIDIA H200, with FlashAttention 3 for target and draft attention
  • Speculation block size: 16 (15 draft tokens per verification step)
  • Sampling: Muse's officially recommended parameters (temperature 1.0, top-p 0.95, top-k 64), with high reasoning strength
  • Maximum new tokens: 4096
  • Prompts: benchmark formatting from z-lab/dflash

We compare autoregressive decoding, the official DFlash drafter (meta-models/Muse-Glimmer-30B-assistant), a community DSpark drafter (DaoCloud/Muse-Glimmer-30B-DSpark), and DFlash 2. All speculative methods propose fifteen draft tokens per verification step.

Acceptance Length

Acceptance length is the per-request mean of completion tokens divided by verification steps. Higher is better.

Task Official DFlash DSpark DFlash 2
GSM8K 5.43 5.45 6.57
MATH-500 5.39 5.01 6.56
HumanEval 4.11 4.33 5.66
MBPP 3.74 4.02 5.30
MT-Bench 3.52 3.59 4.42

Throughput

Throughput is total output tokens divided by end-to-end wall time. Each cell shows output tok/s (speedup vs. autoregressive).

Concurrency 1

Task Autoregressive Official DFlash DSpark DFlash 2
GSM8K 63.9 247.8 (3.88ร—) 236.5 (3.70ร—) 293.7 (4.59ร—)
MATH-500 64.0 246.3 (3.85ร—) 218.4 (3.41ร—) 295.5 (4.62ร—)
HumanEval 65.1 210.5 (3.23ร—) 201.4 (3.09ร—) 266.2 (4.09ร—)
MBPP 63.9 196.8 (3.08ร—) 192.7 (3.02ร—) 264.8 (4.14ร—)
MT-Bench 64.0 164.6 (2.57ร—) 159.7 (2.49ร—) 197.4 (3.08ร—)

Concurrency 8

Task Autoregressive Official DFlash DSpark DFlash 2
GSM8K 476.6 1,574.4 (3.30ร—) 1,456.1 (3.06ร—) 1,816.6 (3.81ร—)
MATH-500 466.0 1,582.9 (3.40ร—) 1,386.0 (2.97ร—) 1,859.3 (3.99ร—)
HumanEval 499.9 1,419.8 (2.84ร—) 1,315.4 (2.63ร—) 1,784.9 (3.57ร—)
MBPP 491.6 1,278.9 (2.60ร—) 1,266.6 (2.58ร—) 1,719.7 (3.50ร—)
MT-Bench 470.0 1,078.4 (2.29ร—) 1,052.6 (2.24ร—) 1,288.9 (2.74ร—)

Concurrency 32

Task Autoregressive Official DFlash DSpark DFlash 2
GSM8K 1,705.6 2,330.3 (1.37ร—) 2,301.7 (1.35ร—) 2,818.3 (1.65ร—)
MATH-500 1,710.2 2,427.3 (1.42ร—) 2,185.0 (1.28ร—) 2,869.6 (1.68ร—)
HumanEval 1,798.1 2,170.0 (1.21ร—) 2,068.9 (1.15ร—) 2,780.2 (1.55ร—)
MBPP 1,717.5 1,964.6 (1.14ร—) 2,006.5 (1.17ร—) 2,685.4 (1.56ร—)
MT-Bench 1,721.8 1,668.0 (0.97ร—) 1,627.2 (0.95ร—) 1,975.5 (1.15ร—)

Citation

If you find DFlash 2 useful, please cite:

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}

Please also cite the original DFlash paper:

@inproceedings{chen2026dflash,
  title     = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author    = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  booktitle = {International Conference on Machine Learning (ICML)},
  year      = {2026}
}
Downloads last month
5,605
Safetensors
Model size
3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for z-lab/Muse-Glimmer-30B-DFlash2

Finetuned
(49)
this model

Space using z-lab/Muse-Glimmer-30B-DFlash2 1

Collection including z-lab/Muse-Glimmer-30B-DFlash2