Qwen3.8-27B-GLA-g2-FP8

FP8 build of Qwen3.8-27B-GLA-g2, the training-free Grouped Latent Attention with 2 groups (GLA-g2) retrofit of Qwen3.8-27B: 31.5 GB instead of 55.7 GB, with the same 2× smaller KV cache. On SWE-bench Verified it resolved 0.403 over 3 runs, against 0.395 for the bf16 base model and 0.399 for Qwen's own FP8 build.

Serving: this checkpoint uses a custom attention architecture. It runs in vLLM with the qwen-mla-vllm plugin (experimental; see Serving). It does not load in transformers or stock vLLM. Includes the base model's vision tower and MTP head.

Qwen3.8-27B (bf16) Qwen/Qwen3.8-27B-FP8 Qwen3.8-27B-GLA-g2 (bf16) Qwen3.8-27B-GLA-g2-FP8
Weights 55.6 GB 30.9 GB 55.7 GB 31.5 GB
KV cache per token per GPU, bf16 (TP=2) 32 KiB 32 KiB 16 KiB 16 KiB
SWE-bench Verified (runs) 0.395 (4) 0.399 (3) 0.394 (3) 0.403 (3)
vs bf16 base (paired, 95% CI) — +0.5 [−1.5, +2.5] −0.1 [−2.1, +2.0] +0.9 [−1.2, +2.9]
Output tokens on SWE-bench 1.00× 1.00× 0.86× 0.82×

All four include the vision tower and the MTP head.

Serving (vLLM plugin, experimental)

pip install "git+https://github.com/sootaugur/qwen-mla-vllm"      # installs vllm==0.27.1
vllm serve TelperionAI/Qwen3.8-27B-GLA-g2-FP8 --tensor-parallel-size 2 --reasoning-parser qwen3 \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'

--speculative-config turns on MTP speculative decoding (optional*). Images work as with the base model; for text-only serving add --limit-mm-per-prompt '{"image": 0, "video": 0}'.

On PCIe-only GPUs (no NVLink, e.g. RTX PRO or GeForce), add --disable-custom-all-reduce at TP=2.

The plugin is experimental. This version has been verified on RTX PRO 6000 (Blackwell) at TP=1 and TP=2, with vLLM 0.27.1 only, and has seen little real-world use so far. It needs a GPU that vLLM runs block-FP8 on (Hopper or Blackwell; only Blackwell has been verified). There vLLM selects CutlassFp8BlockScaledMMKernel. The first start JIT-compiles the fast decode kernel, which needs nvcc and ninja. Without them the plugin falls back to Triton kernels: correct, but slower. See the plugin README for requirements, configuration and verification.

Speed (MLA-FP8, TP=1, RTX PRO 6000, default FP8 KV): 43 / 225 tok/s at 1 / 8 concurrent requests without MTP, 54 / 294 with MTP k=3*. A full benchmark has not been run yet. For bf16 numbers, see the parent model.

*MTP measurements are from one GPU type (RTX PRO 6000, very high memory bandwidth). The benefit depends on hardware: where weight reads dominate decode (lower-bandwidth GPUs) it is likely larger. With an FP8 KV cache on that GPU, k=3 gave 1.6x (bf16 build) and 1.3x (FP8 build) at 1 concurrent request; measure on your own hardware.

What this is

The model is Qwen3.8-27B-GLA-g2: see its card for the method, calibration data and full results. Only the storage format of the weights changes here.

  • Format: compressed-tensors. Weights are FP8 e4m3 with 128×128 block scales. Activations are quantized to FP8 dynamically per token, in groups of 128.
  • Kept in bf16: the latent projections (kv_a_proj, k_up_proj, v_up_proj, k_rope_proj, and the kv_b_proj the plugin builds from them at load time), the linear-attention in_proj_a / in_proj_b and norms, the embeddings and lm_head, the vision tower and the MTP head (both the base model's, unchanged).
  • KV cache: FP8 by default, with per-layer scales calibrated on this model (config mla_kv_cache_scales); --kv-cache-dtype bfloat16 switches to bf16. The SWE-bench results below were measured with a bf16 KV cache; FP8 KV measured within run-to-run noise (see The family).

How it was built. The retrofit differs from Qwen3.8-27B only in its 16 full-attention layers, so it was grafted onto TelperionAI/Qwen3.8-27B-FP8-block-AWQ, an FP8 block build of the base model with AWQ smoothing of the MLPs:

  • MLPs, linear-attention layers, o_proj, embeddings and norms: taken unchanged from that build. The AWQ smoothing touches only the MLP mappings, not the attention inputs, so no attention tensor needed rescaling.
  • q_proj (a fixed permutation of the base q_proj under partial RoPE): re-quantized from the retrofit's bf16 weights with the recipe's own calibration-free FP8 step (block 128×128, amax/448). Applied to the base weights, that step reproduces the published FP8 q_proj / k_proj / o_proj bit for bit.
  • Latent projections and attention norms: the retrofit's bf16 tensors.

Results

All comparisons are paired: same prompts, same harness, same serving settings, against the same bf16 base-model runs as the parent card.

SWE-bench Verified

evalscope, BM25-retrieved context (princeton-nlp/SWE-bench_bm25_40K), thinking on, temperature 1.0, 500 instances.

runs resolved vs bf16 base (paired, 95% CI) output tokens
Qwen3.8-27B (bf16) 4 0.395 — 1.00×
Qwen/Qwen3.8-27B-FP8 3 0.399 +0.5 [−1.5, +2.5] 1.00×
Qwen3.8-27B-GLA-g2 (bf16) 3 0.394 −0.1 [−2.1, +2.0] 0.86×
Qwen3.8-27B-GLA-g2-FP8 3 0.403 (204 / 199 / 202) +0.9 [−1.2, +2.9] 0.82×

Against its own bf16 parent (paired, 3 runs each), the FP8 build differs by +0.9 [−1.2, +3.0] points, which is not significant. Results cover SWE-bench Verified only. HELMET was not re-run on the FP8 build. Vision spot check (100 ChartQA / 100 DocVQA items, with MTP on): 88.0 / 96.1; see the parent's card for the paired vision comparison.

Fidelity to the base model

Teacher-forced NLL on text the bf16 base model generated, excess over the bf16 base (lower is closer):

agentic coding traces SWE-bench Verified conversations
Qwen/Qwen3.8-27B-FP8 +0.003 +0.006
Qwen3.8-27B-GLA-g2 (bf16) +0.045 +0.050
Qwen3.8-27B-GLA-g2-FP8 +0.050 +0.056

FP8 adds about the same error to the retrofit (+0.004 to +0.006) as it adds to the base model. The two sources of error simply add.

Generation check: in thinking mode on SWE-bench prompts, all generations stopped normally, patches were present, and lengths were within sampling variation of the base model.

KV cache and tensor parallelism

Identical to the parent; see its card. This model is meant for TP=1 and TP=2 serving.

The family

Two retrofits of Qwen3.8-27B, each in four weight formats. All keep the vision tower and MTP head, and all ship calibrated FP8 KV-cache scales.

weights MLA (1 GPU) GLA-g2 (2 GPUs) size KV cache default
bf16 Qwen3.8-27B-MLA Qwen3.8-27B-GLA-g2 55.7 GB bf16 (FP8 opt-in)
FP8 -MLA-FP8 -GLA-g2-FP8 31.5 GB FP8
INT4 (AWQ+GPTQ) -MLA-INT4 -GLA-g2-INT4 22.9 GB FP8
EXL3 4.0 bpw -MLA-EXL3-4.0bpw -GLA-g2-EXL3-4.0bpw 17.2 GB FP8

Fidelity to the bf16 base model, teacher-forced on 562k held-out positions: top-1 disagreement where the base model is confident (top-1 vs top-2 logprob margin 2-5; lower is closer). Quantized builds measured with their default FP8 KV cache.

bf16 FP8 INT4 EXL3 4.0 bpw
base model's own quantized builds - 0.76% 0.59% 0.59%
MLA 1.09% 1.15% 1.26% 1.56%
GLA-g2 2.47% 3.16% 2.86% 3.62%

FP8 KV on the bf16 retrofits measured identical (1.09% / 2.47%); its effect is within run-to-run noise.

Capacity, full 262k-token sequences that fit at once (RTX PRO 6000, 96 GB, default settings with vision on, no MTP). Base model, bf16: 1.9 on one GPU.

bf16 bf16 + FP8 KV FP8 INT4 EXL3
MLA, 1 GPU 3.7 7.3 12.7 14.5 15.6
GLA-g2, 2 GPUs (TP=2) 13.7 26.9 31.7 33.9 34.9

Limitations

  • Serving requires the experimental vLLM plugin (vLLM 0.27.1 only, for now).
  • Only a spot check of speed for this build (see Serving).
  • Evaluated on SWE-bench Verified and teacher-forced fidelity. Other domains and long-context benchmarks were not re-run for the FP8 build.
  • Results are from one harness and sampling configuration. Absolute scores differ across harnesses, so the paired differences are the meaningful numbers.

License and attribution

Released under Apache-2.0, the license of the base model. Derived from Qwen/Qwen3.8-27B by the Qwen team. This is an independent project, not affiliated with or endorsed by the Qwen team.

Citation

@misc{telperion2026qwen38glafp8,
  title  = {Qwen3.8-27B-GLA-g2-FP8: FP8 build of a training-free latent attention retrofit},
  author = {TelperionAI},
  year   = {2026},
  url    = {https://huggingface.co/TelperionAI/Qwen3.8-27B-GLA-g2-FP8}
}
Downloads last month
38
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TelperionAI/Qwen3.8-27B-GLA-g2-FP8

Base model

Qwen/Qwen3.8-27B
Quantized
(3)
this model

Collection including TelperionAI/Qwen3.8-27B-GLA-g2-FP8