mistral-7b-instruct-v0.3-q8 (MLX, CBA artifact)

MLX-format 8-bit (Q8) variant of mistralai/Mistral-7B-Instruct-v0.3.

This is one of the 15 model artifacts from the paper:

Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels Plawan Kumar Rath, Rahul Maliakkal. IEEE Cloud Summit 2026. Code: https://github.com/plawanrath/compression-bias-amplification arXiv: https://arxiv.org/abs/2605.15208

Quantization

Weight-only post-training quantization via mlx_lm.convert:

  • bits: 8
  • group_size: 64
  • mode: affine

How this artifact was produced

python -m mlx_lm.convert \
    --hf-path mistralai/Mistral-7B-Instruct-v0.3 \
    --mlx-path ./mistral-7b-instruct-v0.3-q8 \
    --quantize \
    --q-bits 8 \
    --q-group-size 64

This is the exact artifact used to produce the inference results in §4.3 of the paper (911,100 records over BBQ ambiguous, 5 seeds × 12,148 items × 15 configs).

Usage (MLX)

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("plawanrath/mistral-7b-instruct-v0.3-q8-mlx-cba")
prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Hello!"}],
    add_generation_prompt=True,
    tokenize=False,
)
print(generate(model, tokenizer, prompt=prompt, max_tokens=128))

Or via CLI:

mlx_lm.generate --model plawanrath/mistral-7b-instruct-v0.3-q8-mlx-cba --prompt "Hello!"

Paper findings relevant to this variant

The paper documents a dose-response relationship between quantization aggressiveness and emergent stereotypical behavior on BBQ ambiguous questions:

Variant % of BF16-unbiased items that became biased
Q8 0.1–0.9%
Q6 0.3–1.3%
Q4 2.2–5.6%
Q3 6.0–21.1%

These changes are largely invisible to perplexity (<0.5% shift at Q8, <3% at Q4 across all three families). Treat any deployment of compressed instruction-tuned models on fairness-sensitive tasks accordingly.

Model details

License

Inherited from the base model (apache-2.0). See the upstream model page for the full license text.

Citation

@inproceedings{rath2026quantization,
  title     = { Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels },
  author    = {Rath, Plawan Kumar and Maliakkal, Rahul},
  booktitle = { IEEE Cloud Summit 2026 },
  year      = {2026},
  eprint    = {2605.15208},
  archivePrefix = {arXiv},
  url       = {https://arxiv.org/abs/2605.15208}
}

Disclaimer

Plawan Kumar Rath (@plawanrath). This work was conducted in the author's personal capacity. The views expressed here are those of the author and do not necessarily reflect the views of Meta.

Downloads last month
31
Safetensors
Model size
7B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for plawanrath/mistral-7b-instruct-v0.3-q8-mlx-cba

Quantized
(299)
this model

Collection including plawanrath/mistral-7b-instruct-v0.3-q8-mlx-cba

Paper for plawanrath/mistral-7b-instruct-v0.3-q8-mlx-cba