Supra2-IMG-ONNX
Browser/WebGPU-ready export of SupraLabs/Supra2-IMG (a ~104M-parameter rectified-flow DiT: Flan-T5-Base context + SD VAE). The original ships as a raw PyTorch checkpoint with custom CUDA code, which browsers cannot run — this repo contains the same weights exported to ONNX so any web app can download them (once, then cached) and run fully on-device via ONNX Runtime WebGPU.
Files
| Path | Size | What |
|---|---|---|
dit/model.onnx |
398MB | DiT denoiser, fp32 (fp16 was measurably lossy for this bf16-trained net) |
vae_decoder/model.onnx |
189MB | SD VAE decoder (fp32, bit-exact export) |
text_encoder/encoder_model.onnx |
419MB | Flan-T5 encoder, fp32 (snapshotted from Xenova/flan-t5-base). Deliberately fp32: fp16 silently NaNs on some GPU/driver combos |
tokenizer.json, tokenizer_config.json, spiece.model, special_tokens_map.json |
~4MB | T5 tokenizer at repo root, for AutoTokenizer.from_pretrained(repo) |
pipeline_config.json |
1KB | Dims, dtypes, VAE scale, default steps/CFG — read by the runtime |
How it was converted
Environment: Python 3.12, torch (CPU), diffusers, transformers, huggingface_hub, onnx, onnxruntime.
pip install torch diffusers transformers huggingface_hub onnx onnxruntime
python tools/convert_supra.py --repo YOURNAME/Supra2-IMG-ONNX
The converter (tools/supra_arch.py + tools/convert_supra.py):
- Downloads
model_final_ema.ptfrom the base repo and loads it strict into a re-implemented SupraDiT (identical parameter names/shapes; the only change is explicit-softmax attention instead of SDPA, which does not export to ONNX). - Exports the DiT to ONNX (opset 20 for GELU-tanh, dynamic batch for CFG, fp32).
- Exports the decode-only half of
stabilityai/sd-vae-ft-msevia diffusers. - Snapshots the fp32 T5 encoder + tokenizer from
Xenova/flan-t5-base(no conversion needed). - Verifies numeric parity in ONNX Runtime (DiT within backend-noise bounds — the net is chaotic, a 1e-7 input perturbation shifts outputs by ~0.2, so exact match is impossible on any backend; VAE cosine = 1.000000) and uploads the folder.
Validated end to end: 20–50-step Euler CFG sampling with real T5 embeddings renders prompt-faithful 256×256 images.
Run it in JavaScript (no build step)
Requirements: a WebGPU browser (Chrome/Edge 113+) and serving over http(s) — WebGPU,
Cache Storage, and ES modules need a secure context, so file:// will not work.
Step 1. Download example_web.js into an empty folder.
Step 2. Create index.html next to it:
<!doctype html>
<html>
<body>
<input id="prompt" value="a lighthouse above violet clouds at dusk" size="50">
<button id="go">Generate</button>
<progress id="bar" max="100" value="0" style="width:420px"></progress>
<span id="line">Idle</span><br>
<canvas id="out"></canvas>
<script src="https://cdn.jsdelivr.net/npm/onnxruntime-web@1.30.0/dist/ort.all.min.js"></script>
<!-- transformers.js is ESM-only (no UMD build): example_web.js imports it from the CDN itself. -->
<script type="module">
import { loadSupra, generate, paint } from './example_web.js';
const REPO = 'Bartholomheow/Supra2-IMG-ONNX';
const bar = document.querySelector('#bar');
const line = document.querySelector('#line');
// Aggregate the three parallel downloads into one stable line, repainted
// at most ~7x/sec — raw per-chunk events would flicker dozens of times/sec.
const seen = {};
let lastPaint = 0;
const mb = (n) => `${(n / 1048576).toFixed(0)}MB`;
const model = await loadSupra(REPO, { onProgress: (p) => {
seen[p.file] = p;
if (performance.now() - lastPaint < 150) return;
lastPaint = performance.now();
let loaded = 0, total = 0, known = true;
for (const f of Object.values(seen)) {
loaded += f.loaded;
if (f.total == null) known = false; else total += f.total;
}
if (known && total > 0) {
line.textContent = `Downloading models… ${mb(loaded)} / ${mb(total)}`;
bar.value = (loaded / total) * 100;
} else {
line.textContent = `Downloading ${p.file}… ${mb(loaded)}`;
bar.removeAttribute('value');
}
}});
line.textContent = 'Model ready';
document.querySelector('#go').onclick = async () => {
line.textContent = 'Generating…';
const { pixels, size } = await generate(model, document.querySelector('#prompt').value);
paint(pixels, size, document.querySelector('#out'));
line.textContent = 'Done';
};
</script>
</body>
</html>
Step 3. Serve and open it (first visit downloads ~800MB once, then it's cached):
npx serve . # or: python3 -m http.server
# open http://localhost:3000 (or :8000)
Run it in TypeScript (npm project)
Step 1. Install the two runtimes:
npm i onnxruntime-web @huggingface/transformers
Step 2. Make sure tsconfig.json targets modern ESM with DOM (Vite templates already do):
{ "compilerOptions": { "target": "es2022", "module": "nodenext", "moduleResolution": "nodenext", "lib": ["es2022", "dom"] } }
Step 3. Copy example_web.ts into your src/ and use it:
import { loadSupra, generate, paint } from './example_web';
const model = await loadSupra(
'Bartholomheow/Supra2-IMG-ONNX',
(p) => console.log(p.file, (p.loaded / 1048576).toFixed(1) + 'MB'),
);
const { pixels, size } = await generate(model, 'a lighthouse above violet clouds at dusk', { seed: 1 });
paint(pixels, size, document.querySelector('canvas')!);
generate also accepts { steps, cfg } (defaults: 50 steps / CFG 3.0, matching the official
repo) and onProgress, which reports { phase: 'encode' | 'denoise' | 'decode', step, steps } —
wire both to a status line and a <progress> bar so long runs stay visible. On CPU without
hardware acceleration the same 50 steps take several minutes; the step counter shows it's alive.
What happens under the hood (both examples)
pipeline_config.jsonis fetched; the three.onnxfiles download with progress into Cache Storage (supra2-img-v1) — usecaches.delete('supra2-img-v1')to force re-download.- The prompt is tokenized to 128 tokens (
AutoTokenizer.from_pretrained(repo)). - The T5 encoder runs (int64 ids + mask), then the DiT Euler-integrates from seeded noise,
z += dt·(v_uncond + cfg·(v_cond − v_uncond)), t = 0→1. z / 0.18215is decoded by the VAE to a 256×256 image. All sessions use the WebGPU provider.
Troubleshooting
No available backend found/ adapter errors → WebGPU needs hardware acceleration enabled in the browser (Chrome/Edge: Settings → System → "Use hardware acceleration"; VMs, remote desktops, and old drivers often lack it). Note: onnxruntime silently falls back to WASM when WebGPU init fails, so the examples verifynavigator.gpu.requestAdapter()themselves instead of trusting the requested backend.- No WebGPU at all? Both examples accept a backend list with a slow CPU fallback:
JS:
loadSupra(REPO, { backends: ['webgpu', 'wasm'] }), TS:loadSupra(REPO, undefined, ['webgpu', 'wasm']). Expect minutes per image on CPU. The loaded backend is returned asmodel.backend. - Blank page / module errors → you opened it via
file://; serve overhttp://localhostinstead. - First load stalls → it's ~1GB; watch the progress callback. Weak iGPUs still work, just slower.
- Black image with no error → almost always NaN poisoning (NaN paints as black). This repo is
fp32 throughout precisely because fp16 silently NaN'd on some GPU/driver combos during testing.
Remaining NaN failures throw loud (
text encoder returned NaN on this backend…) instead of rendering black, and the demo prints pixel stats (range […]vsN NaN values) after every generation to tell dark-but-valid apart from failure. - Suspect a poisoned download cache (truncated file cached as complete)? The demo footer shows the running build id, and a Clear cached models button wipes Cache Storage and reloads. Downloads are now length-validated, and cached entries are re-checked against the server.
- Different repo id → pass it as the first arg:
loadSupra('you/Supra2-IMG-ONNX', …).
Example code index
| File | What |
|---|---|
example_inference.py |
Standalone Python (onnxruntime + transformers, no torch): python example_inference.py --prompt "..." --out out.png |
example_web.js |
The plain-JS version used above (verified with node --check) |
example_web.ts |
The TypeScript version used above (verified with tsc --noEmit against the real packages) |
Limitations
- ~800MB first download (cached afterwards); needs a WebGPU browser (Chrome/Edge 113+).
- 256×256 output, English prompts; default 50 steps / CFG 3.0 (same as the official repo).
License & credit
Apache-2.0, same as the base model. All weights by SupraLabs; conversion only changes the container format.
Model tree for Bartholomheow/Supra2-IMG-ONNX
Base model
SupraLabs/Supra2-IMG