File size: 5,480 Bytes
27e9aaa
 
 
9ff986b
 
 
 
 
 
ef1adab
 
9ff986b
 
 
 
 
 
 
 
 
 
 
 
 
ef1adab
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
---
base_model:
- Cloudflare/clef-flash
library_name: transformers
tags:
- text-generation-inference
- llm-compressor
- vllm
- clef
- fp8
- w8a8
- cloudflare
- systemone
- qwen3.5
- post-train
- image-text-to-typed-output
- multimodal
- structured-output
- classification
- custom-code
license: apache-2.0
language:
- en
pipeline_tag: image-text-to-text
---

# **clef-flash-FP8**

FP8 (W8A8, dynamic) quantization of [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash),
a 9B multimodal model that turns a state and a schema of typed questions into decisions.
Clef-Flash reads text, JSON, images, or video and returns a probability for every allowed option of
every question in a single forward pass, with no free-form generation and no output parsing. This repo quantizes only the backbone's linear layers. The vision encoder, embeddings, `lm_head`,
and linear-attention layers are left in their original precision. For model behavior, input
format, and the Jev/SystemOne API, see the
[original Clef-Flash card](https://huggingface.co/Cloudflare/clef-flash).

## Quantization

| | |
|---|---|
| **Modality** | Image-Text-to-Text |
| **Quantization scheme** | FP8_DYNAMIC (W8A8) |
| **Weights** | FP8, per-channel |
| **Activations** | FP8, per-token, dynamic |
| **Calibration data** | Not required |
| **Format** | compressed-tensors (safetensors) |
| **Tooling** | [LLM Compressor](https://github.com/vllm-project/llm-compressor) |
| **License** | Apache 2.0 |

| Setting | Value |
|---|---|
| **targets** | `Linear` |
| **ignore** | `lm_head`, `embed_tokens`, `visual`, `model.visual`, `linear_attn` |
| **scheme** | `FP8_DYNAMIC` |
| **bypass_divisibility_checks** | `false` |
| **requires_calibration_data** | `false` |

### recipe.yaml

```yaml
default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
        're:.*linear_attn.*']
      scheme: FP8_DYNAMIC
      bypass_divisibility_checks: false
      requires_calibration_data: false
```

Because the scheme is `FP8_DYNAMIC`, weight scales are computed directly from the weights and
activation scales are computed per token at runtime. No calibration dataset is needed.

## Usage

Install `compressed-tensors` alongside `transformers` so the FP8 checkpoint can be loaded:

```bash
pip install torch transformers compressed-tensors pillow
```

Usage is the same as for Clef-Flash:

```python
import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("prithivMLmods/clef-flash-FP8")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))
```

The `systemone(model, processor, request)` helper and image/video inputs work as described in the
Clef-Flash card.

**Hardware:** FP8 W8A8 compute needs a GPU with FP8 support (Ada, Hopper, or newer). On older
GPUs, FP8 weights may only give memory savings, depending on the runtime.

## Reproducing the quantization

```python
from transformers import AutoModelForImageTextToText, AutoProcessor
from llmcompressor import oneshot

src = "Cloudflare/clef-flash"   # local path from snapshot_download works too
dst = "clef-flash-FP8"

model = AutoModelForImageTextToText.from_pretrained(src, torch_dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(src)

oneshot(model=model, recipe="recipe.yaml")   # data-free, no dataset argument

model.save_pretrained(dst, save_compressed=True)
processor.save_pretrained(dst)
```

Then copy these files from the original Clef-Flash repo into `clef-flash-FP8/` unchanged:
`joint_head.safetensors`, `joint_head_config.json`, `joint_schema_model.py`.

## Notes and limitations

- **Joint head:** the joint schema head is stored separately and is kept in its original
  precision, because the recipe quantizes only the backbone. If you modify `load_release_model` or
  re-export the model, make sure the head is not quantized.
- **Excluded modules:** `linear_attn`, the vision encoder, embeddings, and `lm_head` stay
  unquantized, so the size reduction is somewhat smaller than a full 2x versus BF16.
- **Loading path:** this checkpoint is intended for the custom `joint_schema_model.py` loader.
  General-purpose serving engines will not run the joint head.

## License

Apache-2.0, following [Cloudflare/clef-flash](https://huggingface.co/Cloudflare/clef-flash) and the
base model [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B).