Qwen3.8-27B-GLA-g2-FP8
FP8 build of Qwen3.8-27B-GLA-g2, the training-free Grouped Latent Attention with 2 groups (GLA-g2) retrofit of Qwen3.8-27B: 31.5 GB instead of 55.7 GB, with the same 2× smaller KV cache. On SWE-bench Verified it resolved 0.403 over 3 runs, against 0.395 for the bf16 base model and 0.399 for Qwen's own FP8 build.
Serving: this checkpoint uses a custom attention architecture. It runs in vLLM with the qwen-mla-vllm plugin (experimental; see Serving). It does not load in
transformersor stock vLLM. Includes the base model's vision tower and MTP head.
| Qwen3.8-27B (bf16) | Qwen/Qwen3.8-27B-FP8 | Qwen3.8-27B-GLA-g2 (bf16) | Qwen3.8-27B-GLA-g2-FP8 | |
|---|---|---|---|---|
| Weights | 55.6 GB | 30.9 GB | 55.7 GB | 31.5 GB |
| KV cache per token per GPU, bf16 (TP=2) | 32 KiB | 32 KiB | 16 KiB | 16 KiB |
| SWE-bench Verified (runs) | 0.395 (4) | 0.399 (3) | 0.394 (3) | 0.403 (3) |
| vs bf16 base (paired, 95% CI) | — | +0.5 [−1.5, +2.5] | −0.1 [−2.1, +2.0] | +0.9 [−1.2, +2.9] |
| Output tokens on SWE-bench | 1.00× | 1.00× | 0.86× | 0.82× |
All four include the vision tower and the MTP head.
Serving (vLLM plugin, experimental)
pip install "git+https://github.com/sootaugur/qwen-mla-vllm" # installs vllm==0.27.1
vllm serve TelperionAI/Qwen3.8-27B-GLA-g2-FP8 --tensor-parallel-size 2 --reasoning-parser qwen3 \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'
--speculative-config turns on MTP speculative decoding (optional*). Images work as with the
base model; for text-only serving add --limit-mm-per-prompt '{"image": 0, "video": 0}'.
On PCIe-only GPUs (no NVLink, e.g. RTX PRO or GeForce), add --disable-custom-all-reduce at TP=2.
The plugin is experimental. This version has been verified on RTX PRO 6000 (Blackwell) at
TP=1 and TP=2, with vLLM 0.27.1 only, and has seen little real-world use so far. It needs a GPU
that vLLM runs block-FP8 on (Hopper or Blackwell; only Blackwell has been verified). There vLLM selects
CutlassFp8BlockScaledMMKernel. The first start JIT-compiles the fast decode kernel, which needs
nvcc and ninja. Without them the plugin falls back to Triton kernels: correct, but slower.
See the plugin README for requirements, configuration and verification.
Speed (MLA-FP8, TP=1, RTX PRO 6000, default FP8 KV): 43 / 225 tok/s at 1 / 8 concurrent requests without MTP, 54 / 294 with MTP k=3*. A full benchmark has not been run yet. For bf16 numbers, see the parent model.
*MTP measurements are from one GPU type (RTX PRO 6000, very high memory bandwidth). The benefit depends on hardware: where weight reads dominate decode (lower-bandwidth GPUs) it is likely larger. With an FP8 KV cache on that GPU, k=3 gave 1.6x (bf16 build) and 1.3x (FP8 build) at 1 concurrent request; measure on your own hardware.
What this is
The model is Qwen3.8-27B-GLA-g2: see its card for the method, calibration data and full results. Only the storage format of the weights changes here.
- Format: compressed-tensors. Weights are FP8 e4m3 with 128×128 block scales. Activations are quantized to FP8 dynamically per token, in groups of 128.
- Kept in bf16: the latent projections (
kv_a_proj,k_up_proj,v_up_proj,k_rope_proj, and thekv_b_projthe plugin builds from them at load time), the linear-attentionin_proj_a/in_proj_band norms, the embeddings andlm_head, the vision tower and the MTP head (both the base model's, unchanged). - KV cache: FP8 by default, with per-layer scales calibrated on this model (config
mla_kv_cache_scales);--kv-cache-dtype bfloat16switches to bf16. The SWE-bench results below were measured with a bf16 KV cache; FP8 KV measured within run-to-run noise (see The family).
How it was built. The retrofit differs from Qwen3.8-27B only in its 16 full-attention layers, so it was grafted onto TelperionAI/Qwen3.8-27B-FP8-block-AWQ, an FP8 block build of the base model with AWQ smoothing of the MLPs:
- MLPs, linear-attention layers,
o_proj, embeddings and norms: taken unchanged from that build. The AWQ smoothing touches only the MLP mappings, not the attention inputs, so no attention tensor needed rescaling. q_proj(a fixed permutation of the baseq_projunder partial RoPE): re-quantized from the retrofit's bf16 weights with the recipe's own calibration-free FP8 step (block 128×128, amax/448). Applied to the base weights, that step reproduces the published FP8q_proj/k_proj/o_projbit for bit.- Latent projections and attention norms: the retrofit's bf16 tensors.
Results
All comparisons are paired: same prompts, same harness, same serving settings, against the same bf16 base-model runs as the parent card.
SWE-bench Verified
evalscope, BM25-retrieved context (princeton-nlp/SWE-bench_bm25_40K), thinking on,
temperature 1.0, 500 instances.
| runs | resolved | vs bf16 base (paired, 95% CI) | output tokens | |
|---|---|---|---|---|
| Qwen3.8-27B (bf16) | 4 | 0.395 | — | 1.00× |
| Qwen/Qwen3.8-27B-FP8 | 3 | 0.399 | +0.5 [−1.5, +2.5] | 1.00× |
| Qwen3.8-27B-GLA-g2 (bf16) | 3 | 0.394 | −0.1 [−2.1, +2.0] | 0.86× |
| Qwen3.8-27B-GLA-g2-FP8 | 3 | 0.403 (204 / 199 / 202) | +0.9 [−1.2, +2.9] | 0.82× |
Against its own bf16 parent (paired, 3 runs each), the FP8 build differs by +0.9 [−1.2, +3.0] points, which is not significant. Results cover SWE-bench Verified only. HELMET was not re-run on the FP8 build. Vision spot check (100 ChartQA / 100 DocVQA items, with MTP on): 88.0 / 96.1; see the parent's card for the paired vision comparison.
Fidelity to the base model
Teacher-forced NLL on text the bf16 base model generated, excess over the bf16 base (lower is closer):
| agentic coding traces | SWE-bench Verified conversations | |
|---|---|---|
| Qwen/Qwen3.8-27B-FP8 | +0.003 | +0.006 |
| Qwen3.8-27B-GLA-g2 (bf16) | +0.045 | +0.050 |
| Qwen3.8-27B-GLA-g2-FP8 | +0.050 | +0.056 |
FP8 adds about the same error to the retrofit (+0.004 to +0.006) as it adds to the base model. The two sources of error simply add.
Generation check: in thinking mode on SWE-bench prompts, all generations stopped normally, patches were present, and lengths were within sampling variation of the base model.
KV cache and tensor parallelism
Identical to the parent; see its card. This model is meant for TP=1 and TP=2 serving.
The family
Two retrofits of Qwen3.8-27B, each in four weight formats. All keep the vision tower and MTP head, and all ship calibrated FP8 KV-cache scales.
| weights | MLA (1 GPU) | GLA-g2 (2 GPUs) | size | KV cache default |
|---|---|---|---|---|
| bf16 | Qwen3.8-27B-MLA | Qwen3.8-27B-GLA-g2 | 55.7 GB | bf16 (FP8 opt-in) |
| FP8 | -MLA-FP8 | -GLA-g2-FP8 | 31.5 GB | FP8 |
| INT4 (AWQ+GPTQ) | -MLA-INT4 | -GLA-g2-INT4 | 22.9 GB | FP8 |
| EXL3 4.0 bpw | -MLA-EXL3-4.0bpw | -GLA-g2-EXL3-4.0bpw | 17.2 GB | FP8 |
Fidelity to the bf16 base model, teacher-forced on 562k held-out positions: top-1 disagreement where the base model is confident (top-1 vs top-2 logprob margin 2-5; lower is closer). Quantized builds measured with their default FP8 KV cache.
| bf16 | FP8 | INT4 | EXL3 4.0 bpw | |
|---|---|---|---|---|
| base model's own quantized builds | - | 0.76% | 0.59% | 0.59% |
| MLA | 1.09% | 1.15% | 1.26% | 1.56% |
| GLA-g2 | 2.47% | 3.16% | 2.86% | 3.62% |
FP8 KV on the bf16 retrofits measured identical (1.09% / 2.47%); its effect is within run-to-run noise.
Capacity, full 262k-token sequences that fit at once (RTX PRO 6000, 96 GB, default settings with vision on, no MTP). Base model, bf16: 1.9 on one GPU.
| bf16 | bf16 + FP8 KV | FP8 | INT4 | EXL3 | |
|---|---|---|---|---|---|
| MLA, 1 GPU | 3.7 | 7.3 | 12.7 | 14.5 | 15.6 |
| GLA-g2, 2 GPUs (TP=2) | 13.7 | 26.9 | 31.7 | 33.9 | 34.9 |
Limitations
- Serving requires the experimental vLLM plugin (vLLM 0.27.1 only, for now).
- Only a spot check of speed for this build (see Serving).
- Evaluated on SWE-bench Verified and teacher-forced fidelity. Other domains and long-context benchmarks were not re-run for the FP8 build.
- Results are from one harness and sampling configuration. Absolute scores differ across harnesses, so the paired differences are the meaningful numbers.
License and attribution
Released under Apache-2.0, the license of the base model. Derived from Qwen/Qwen3.8-27B by the Qwen team. This is an independent project, not affiliated with or endorsed by the Qwen team.
Citation
@misc{telperion2026qwen38glafp8,
title = {Qwen3.8-27B-GLA-g2-FP8: FP8 build of a training-free latent attention retrofit},
author = {TelperionAI},
year = {2026},
url = {https://huggingface.co/TelperionAI/Qwen3.8-27B-GLA-g2-FP8}
}
- Downloads last month
- 38