Supra2-IMG-ONNX

Browser/WebGPU-ready export of SupraLabs/Supra2-IMG (a ~104M-parameter rectified-flow DiT: Flan-T5-Base context + SD VAE). The original ships as a raw PyTorch checkpoint with custom CUDA code, which browsers cannot run — this repo contains the same weights exported to ONNX so any web app can download them (once, then cached) and run fully on-device via ONNX Runtime WebGPU.

Files

Path Size What
dit/model.onnx 398MB DiT denoiser, fp32 (fp16 was measurably lossy for this bf16-trained net)
vae_decoder/model.onnx 189MB SD VAE decoder (fp32, bit-exact export)
text_encoder/encoder_model.onnx 419MB Flan-T5 encoder, fp32 (snapshotted from Xenova/flan-t5-base). Deliberately fp32: fp16 silently NaNs on some GPU/driver combos
tokenizer.json, tokenizer_config.json, spiece.model, special_tokens_map.json ~4MB T5 tokenizer at repo root, for AutoTokenizer.from_pretrained(repo)
pipeline_config.json 1KB Dims, dtypes, VAE scale, default steps/CFG — read by the runtime

How it was converted

Environment: Python 3.12, torch (CPU), diffusers, transformers, huggingface_hub, onnx, onnxruntime.

pip install torch diffusers transformers huggingface_hub onnx onnxruntime
python tools/convert_supra.py --repo YOURNAME/Supra2-IMG-ONNX

The converter (tools/supra_arch.py + tools/convert_supra.py):

  1. Downloads model_final_ema.pt from the base repo and loads it strict into a re-implemented SupraDiT (identical parameter names/shapes; the only change is explicit-softmax attention instead of SDPA, which does not export to ONNX).
  2. Exports the DiT to ONNX (opset 20 for GELU-tanh, dynamic batch for CFG, fp32).
  3. Exports the decode-only half of stabilityai/sd-vae-ft-mse via diffusers.
  4. Snapshots the fp32 T5 encoder + tokenizer from Xenova/flan-t5-base (no conversion needed).
  5. Verifies numeric parity in ONNX Runtime (DiT within backend-noise bounds — the net is chaotic, a 1e-7 input perturbation shifts outputs by ~0.2, so exact match is impossible on any backend; VAE cosine = 1.000000) and uploads the folder.

Validated end to end: 20–50-step Euler CFG sampling with real T5 embeddings renders prompt-faithful 256×256 images.

Run it in JavaScript (no build step)

Requirements: a WebGPU browser (Chrome/Edge 113+) and serving over http(s) — WebGPU, Cache Storage, and ES modules need a secure context, so file:// will not work.

Step 1. Download example_web.js into an empty folder.

Step 2. Create index.html next to it:

<!doctype html>
<html>
<body>
  <input id="prompt" value="a lighthouse above violet clouds at dusk" size="50">
  <button id="go">Generate</button>
  <progress id="bar" max="100" value="0" style="width:420px"></progress>
  <span id="line">Idle</span><br>
  <canvas id="out"></canvas>
  <script src="https://cdn.jsdelivr.net/npm/onnxruntime-web@1.30.0/dist/ort.all.min.js"></script>
  <!-- transformers.js is ESM-only (no UMD build): example_web.js imports it from the CDN itself. -->
  <script type="module">
    import { loadSupra, generate, paint } from './example_web.js';
    const REPO = 'Bartholomheow/Supra2-IMG-ONNX';
    const bar = document.querySelector('#bar');
    const line = document.querySelector('#line');
    // Aggregate the three parallel downloads into one stable line, repainted
    // at most ~7x/sec — raw per-chunk events would flicker dozens of times/sec.
    const seen = {};
    let lastPaint = 0;
    const mb = (n) => `${(n / 1048576).toFixed(0)}MB`;
    const model = await loadSupra(REPO, { onProgress: (p) => {
      seen[p.file] = p;
      if (performance.now() - lastPaint < 150) return;
      lastPaint = performance.now();
      let loaded = 0, total = 0, known = true;
      for (const f of Object.values(seen)) {
        loaded += f.loaded;
        if (f.total == null) known = false; else total += f.total;
      }
      if (known && total > 0) {
        line.textContent = `Downloading models… ${mb(loaded)} / ${mb(total)}`;
        bar.value = (loaded / total) * 100;
      } else {
        line.textContent = `Downloading ${p.file}… ${mb(loaded)}`;
        bar.removeAttribute('value');
      }
    }});
    line.textContent = 'Model ready';
    document.querySelector('#go').onclick = async () => {
      line.textContent = 'Generating…';
      const { pixels, size } = await generate(model, document.querySelector('#prompt').value);
      paint(pixels, size, document.querySelector('#out'));
      line.textContent = 'Done';
    };
  </script>
</body>
</html>

Step 3. Serve and open it (first visit downloads ~800MB once, then it's cached):

npx serve .        # or: python3 -m http.server
# open http://localhost:3000 (or :8000)

Run it in TypeScript (npm project)

Step 1. Install the two runtimes:

npm i onnxruntime-web @huggingface/transformers

Step 2. Make sure tsconfig.json targets modern ESM with DOM (Vite templates already do):

{ "compilerOptions": { "target": "es2022", "module": "nodenext", "moduleResolution": "nodenext", "lib": ["es2022", "dom"] } }

Step 3. Copy example_web.ts into your src/ and use it:

import { loadSupra, generate, paint } from './example_web';

const model = await loadSupra(
  'Bartholomheow/Supra2-IMG-ONNX',
  (p) => console.log(p.file, (p.loaded / 1048576).toFixed(1) + 'MB'),
);
const { pixels, size } = await generate(model, 'a lighthouse above violet clouds at dusk', { seed: 1 });
paint(pixels, size, document.querySelector('canvas')!);

generate also accepts { steps, cfg } (defaults: 50 steps / CFG 3.0, matching the official repo) and onProgress, which reports { phase: 'encode' | 'denoise' | 'decode', step, steps } — wire both to a status line and a <progress> bar so long runs stay visible. On CPU without hardware acceleration the same 50 steps take several minutes; the step counter shows it's alive.

What happens under the hood (both examples)

  1. pipeline_config.json is fetched; the three .onnx files download with progress into Cache Storage (supra2-img-v1) — use caches.delete('supra2-img-v1') to force re-download.
  2. The prompt is tokenized to 128 tokens (AutoTokenizer.from_pretrained(repo)).
  3. The T5 encoder runs (int64 ids + mask), then the DiT Euler-integrates from seeded noise, z += dt·(v_uncond + cfg·(v_cond − v_uncond)), t = 0→1.
  4. z / 0.18215 is decoded by the VAE to a 256×256 image. All sessions use the WebGPU provider.

Troubleshooting

  • No available backend found / adapter errors → WebGPU needs hardware acceleration enabled in the browser (Chrome/Edge: Settings → System → "Use hardware acceleration"; VMs, remote desktops, and old drivers often lack it). Note: onnxruntime silently falls back to WASM when WebGPU init fails, so the examples verify navigator.gpu.requestAdapter() themselves instead of trusting the requested backend.
  • No WebGPU at all? Both examples accept a backend list with a slow CPU fallback: JS: loadSupra(REPO, { backends: ['webgpu', 'wasm'] }), TS: loadSupra(REPO, undefined, ['webgpu', 'wasm']). Expect minutes per image on CPU. The loaded backend is returned as model.backend.
  • Blank page / module errors → you opened it via file://; serve over http://localhost instead.
  • First load stalls → it's ~1GB; watch the progress callback. Weak iGPUs still work, just slower.
  • Black image with no error → almost always NaN poisoning (NaN paints as black). This repo is fp32 throughout precisely because fp16 silently NaN'd on some GPU/driver combos during testing. Remaining NaN failures throw loud (text encoder returned NaN on this backend…) instead of rendering black, and the demo prints pixel stats (range […] vs N NaN values) after every generation to tell dark-but-valid apart from failure.
  • Suspect a poisoned download cache (truncated file cached as complete)? The demo footer shows the running build id, and a Clear cached models button wipes Cache Storage and reloads. Downloads are now length-validated, and cached entries are re-checked against the server.
  • Different repo id → pass it as the first arg: loadSupra('you/Supra2-IMG-ONNX', …).

Example code index

File What
example_inference.py Standalone Python (onnxruntime + transformers, no torch): python example_inference.py --prompt "..." --out out.png
example_web.js The plain-JS version used above (verified with node --check)
example_web.ts The TypeScript version used above (verified with tsc --noEmit against the real packages)

Limitations

  • ~800MB first download (cached afterwards); needs a WebGPU browser (Chrome/Edge 113+).
  • 256×256 output, English prompts; default 50 steps / CFG 3.0 (same as the official repo).

License & credit

Apache-2.0, same as the base model. All weights by SupraLabs; conversion only changes the container format.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Bartholomheow/Supra2-IMG-ONNX

Quantized
(1)
this model

Space using Bartholomheow/Supra2-IMG-ONNX 1