QSIV
QSIV is a non-autoregressive decision model. You give it a document (code, logs, a policy, a contract, JSON) of up to 65,536 tokens and 1–8 typed questions about it. In a single forward pass it returns a calibrated probability distribution for each question. It does not generate text.
| Question type | You supply | You get |
|---|---|---|
| Noul (yes/no) | a question | P(yes) |
| Choice | a question and 2–255 options | a probability for each option |
| Score | a question and 2–26 ordered levels | a probability for each level, plus the expected level |
Questions are independent. Adding, removing or reordering the questions in a request never changes another question's answer (bitwise identical), so one request can carry several unrelated checks.
- Developed by: the QSIV authors.
- Model type: decision head on a hybrid Gated DeltaNet / attention transformer. 752,393,024 parameters, derived from Qwen3.5-0.8B-Base.
- Weights: FP8 for production (
qsiv_fp8.safetensors) and bf16 (model.safetensors), safetensors only. - Licence: Apache-2.0.
QSIV vs. Doom: the model was faster than the game
We pointed QSIV at a live ViZDoom marine, with no eyes and no text generation, and asked for a fresh decision on every one of Doom's 35 ticks per second.
Then something unexpected happened: the game became the bottleneck.
Rendering a single Doom tic takes the engine about 3.1 ms at 640×480. QSIV decides in about 2.2 ms with a 4-question request and about 3.0 ms with 8 questions. The model was keeping pace with the game engine itself, and on lighter requests it was ahead.
So we had to speed the game up to keep the model fed. We moved the engine onto worker threads and dropped the render to 320×240 (about 1.0 ms per tic). Only then did the bottleneck swing back to QSIV, which, with 4–8 worker threads (each with its own CUDA stream over the same weights), held 16 live games at 35 Hz at once at roughly 650–710 decisions per second.
Click the clip for the full video (62 s). The right-hand panel shows everything the model sees and says: the exact input document, every probability, the keys pressed, and the latency of every single decision on a live graph.
Numbers
Measured on one RTX PRO 6000 Blackwell, with the production FP8 runtime and CUDA-graph replay, one request at a time (batch size 1), 0 output tokens. Latency is the full decide() call (tokenization, forward pass and calibration) over 153,314 live decisions.
| Decision latency p50 / p90 / p99 / max | 2.45 / 3.18 / 3.34 / 6.5 ms |
| 4-question request | about 2.2 ms |
| 8-question request | about 3.0 ms |
| Fresh-document microbenchmark (p50 / p99) | 2.26 / 3.03 ms |
| Decision rate | about 400 decisions/s (the game needs 35/s) |
| Share of each 28.6 ms game tic | about 10% |
| Game engine alone, 640×480 | about 3.1 ms per tic |
| Game engine alone, 320×240 | about 1.0 ms per tic |
| One model, many players | 16 live games at 35 Hz, about 650–710 decisions/s |
| Peak GPU memory | 4.3 GB |
The whole pipeline (decision plus engine step) runs about 5× faster than real time.
At a glance
Measured on one RTX PRO 6000 Blackwell against two other open typed-decision models, Kev-0.8B (the same base model and
size) and Laya Typed-Decisions. Every request carries a new document. Details and all numbers: benchmarks/COMPARISON.md.
Against Jev (TypeSafe's first System One model, available only as a cloud API). We did not measure Jev. Every Jev
latency below was reported by others: by TypeSafe itself, by API routers (OpenRouter, OrcaRouter), by Every, and by the Kev
project, which published its own Jev runs. Those Jev numbers therefore include the network. The QSIV numbers are our own
measurements: svc.decide() inside the application, on a new document given as text (tokenization included), so there is
no network at all. On requests of the Kev project's published Jev suites' size (~400 tokens), Jev's median is 222–232 ms and
QSIV's is 3.0 ms: 74× faster. The fastest QSIV decision is 1.76 ms. Sources and conditions:
benchmarks/JEV_LATENCY.md and benchmarks/jev_latency_sources.json.
A third party measured Jev at five document lengths (1K–32K tokens; jev-1.13 via OpenRouter, warm connection, huggingface.co/datasets/lisonallen/jev-ai-benchmark). At the same lengths QSIV is 10–77× faster: 4.9 ms against Jev's 375 ms at 1K tokens, and 72 ms against 734 ms at 32K.
Self-hosting cost: ~$2 per billion input tokens ($0.002 per million), output free. QSIV is an open-weights model and 100% free to use under the Apache-2.0 licence—there are no API fees or subscription charges whatsoever. To provide an apples-to-apples comparison against commercial APIs like Jev ($42 per billion) and LLMs ($200–10,000 per billion), we measure what QSIV costs to self-host on rented hardware: on one rented RTX PRO 6000 Blackwell ($2.09 per hour on RunPod), the hardware operating cost works out to approximately $2 per billion input tokens across workflows. (Actual self-hosting costs vary across platforms and hardware setups—for example, Vast.ai median $1.51 per hour lowers this by 28%, and owned hardware costs even less). At this self-hosted rate, QSIV is 21× cheaper than Jev on every request, and 100× to 5,000× cheaper than the LLMs' published input prices, before counting the LLMs' output tokens.
| per million input tokens | per billion input tokens | QSIV is cheaper by | |
|---|---|---|---|
| QSIV (self-hosted) | $0.002 | $2 | |
| Jev (published) | $0.042 | $42 | 21× |
| GPT-5.6 Luna (published) | $0.20 | $200 | 100× |
| Claude Haiku 4.5 (published) | $1.00 | $1,000 | 500× |
| GPT-5.6 Terra (published) | $2.00 | $2,000 | 1,000× |
| Claude Fable 5.1 (implied by TypeSafe's 238× claim) | $10.00 | $10,000 | 5,000× |
Per decision: 178–4,389× cheaper than LLMs. TypeSafe compares Jev with LLMs by the cost of each decision on its
workflows, where an LLM also pays for output and reasoning (that is where its "444.6× cheaper" comes from). On the same
published workflows, with QSIV's self-hosting cost applied to the same input, a million decisions cost $19 with QSIV, $390
with Jev and $3,310–120,000 with the LLMs. By TypeSafe's own 444.6× for Jev, QSIV is over 9,000× cheaper than LLMs.
Jev's and the LLMs' costs are published by others, not measured by us. Method, every number and sources:
benchmarks/PRICING.md.
What each open model costs to serve on the same GPU (QSIV, Kev-0.8B and Laya, all as HTTP servers):
Quick start
Requirements: an NVIDIA Blackwell GPU (sm_120, e.g. RTX PRO 6000) with CUDA 13.0, and Python ≥ 3.12. The production runtime is FP8 Triton kernels replayed as CUDA graphs; other GPU generations are not supported by the FP8 path.
Coming soon: support for laptop GPUs, other desktop GPUs, Macs, and phone GPUs. An open-source release of the implementation is also coming soon.
pip install -U huggingface_hub # provides the `hf` command
hf download djinn-san/QSIV --local-dir qsiv && cd qsiv
pip install -e . # torch 2.11, triton 3.6, transformers 5.14, safetensors, numpy
python examples/quickstart.py
from qsiv.runtime.service import QSIVService
svc = QSIVService("qsiv_fp8.safetensors", "calibration.json", tokenizer_dir="tokenizer")
svc.decide({
"state": open("service.log").read(),
"questions": [
{"type": "noul", "text": "Did the deploy at 14:02 fail?"},
{"type": "choice", "text": "Which service failed first?", "options": ["api", "db", "cache"]},
{"type": "score", "text": "How severe is the incident?", "levels": ["none", "low", "medium", "high"]},
],
})
# [{'type': 'noul', 'p_true': ..., 'decision': ..., 'confidence': ...},
# {'type': 'choice', 'choice': k, 'option': ..., 'probs': [...], 'confidence': ...},
# {'type': 'score', 'level': k, 'level_label': ..., 'score': ..., 'uncertainty': ..., 'probs': [...], 'confidence': ...}]
- First run and warm-up (measured on a fresh machine):
- the very first
python examples/quickstart.pytakes about 25 s, because the Triton kernels compile; Triton caches them on disk, and later processes start in about 7 s (model load about 5 s); - inside a process, the first request of each new shape (state bucket × number of questions × question length) takes
about 0.05–0.6 s to build and capture its CUDA-graph plan; every later request of that shape runs at the speeds below.
Plans stay cached within a GPU-memory budget (
QSIVService(..., plan_memory_gb=...), default min(24 GB, 40% of the GPU)): a mixed workload of 12 shapes from 64 to 65K tokens holds about 16 GB; - a server should pre-build the shapes it expects at startup, e.g.
svc.warmup([(512, 1, 32), (512, 1, 64), (4096, 1, 64), (16384, 1, 64), (65536, 1, 64), (512, 8, 64)]), where each tuple is (state bucket, questions, question-block bucket), with buckets listed inconfig.json(a short yes/no question fits the 32-token bucket; a Choice question with its options usually needs 64 or more).
- the very first
- Integrity: verify the downloads with
sha256sum -c qsiv_fp8.safetensors.sha256 model.safetensors.sha256.
Fast paths (all exact: same token ids, same outputs):
- repeated question texts and recently seen long documents are tokenized once and cached;
- longer requests are tokenized in parallel: the document (split only at positions where the tokenizer itself splits) and the questions go to the tokenizer's threads in one call;
- a document you will query repeatedly can be tokenized once with
ids = svc.encode_state(text)and then passed assvc.decide({"state_ids": ids, "questions": [...]}).
bf16 reference path. model.safetensors runs through the PyTorch model definition with
qsiv.reference.ReferenceService (same interface; needs pip install -e ".[reference]"). It is meant for research,
fine-tuning and porting. It is slower than the FP8 runtime, and its decisions match the FP8 runtime's on
about 96% of 422 real held-out questions (mean largest probability difference 0.022; the differences come from FP8
rounding, concentrated on low-margin questions).
Files
| File | Contents |
|---|---|
qsiv_fp8.safetensors |
production weights: FP8 E4M3 matrices with per-tensor scales, static activation scales in the header |
model.safetensors |
the same model in bf16 (GDN gate parameters in fp32) |
calibration.json |
temperatures per question type and per type × request-length band, applied by QSIVService |
config.json, readout.json |
architecture, FP8 contract, request limits, plan buckets; question templates and answer labels |
tokenizer/ |
Qwen3.5 tokenizer |
qsiv/ |
inference package: FP8 Triton runtime, request rendering, model definition |
examples/quickstart.py |
the example above |
benchmarks/ |
BENCHMARKS.md (all results), COMPARISON.md (vs Kev-0.8B and Laya), JEV_LATENCY.md (vs Jev), PRICING.md (cost per token), raw result JSON |
Benchmarks
Speed and document-reading results next to Kev-0.8B and Laya Typed-Decisions: benchmarks/COMPARISON.md.
All numbers are measured on these weights through the production FP8 runtime with the shipped calibration. The baseline is
the untouched base model, run through the same runtime on the same samples. Full tables: benchmarks/BENCHMARKS.md.
Held-out benchmark evaluation: QSIV beats the base model on all 14 evaluation sources (gains from +0.003 to +0.548). Mean accuracy is 0.82, against 0.48 for the base model.
Real code and logs, never seen in training:
| task | base model | QSIV |
|---|---|---|
| parameter binding / direct callers in real Python packages | 0.56 / 0.11 | 0.96 / 0.86 |
| source IP / count of failed SSH logins (Loghub 2.0) | 0.75 / 0.10 | 0.95 / 0.41 |
Long context (synthetic probes; accuracy at 1K / 16K / 65K total tokens):
| task | 1K | 16K | 65K |
|---|---|---|---|
| retrieve a fact from the start of the document | 1.00 | 1.00 | 1.00 |
| apply a later update that overrides an earlier value | 0.98 | 0.65 | 0.85 |
| exact key lookup among near-identical decoys | 0.98 | 0.53 | 0.45 |
| Choice among many options (26 → 255) | 1.00 | 0.72 | 0.30 |
Calibration (expected calibration error, lower is better): 0.010 on independent in-distribution test data and 0.081 on shifted data. By type: yes/no 0.013, Score 0.026, Choice 0.087.
Latency (RTX PRO 6000 Blackwell). Every request carries a new document and new questions, and the time is the
complete svc.decide() call: tokenization, model and post-processing (median, cold L2 cache):
| request tokens | 1 question | 8 questions |
|---|---|---|
| ~60 | 1.76 ms | 2.6 ms |
| ~500 | 3.0 ms | 3.9 ms |
| 4K | 10.8 ms | 12.6 ms |
| 16K | 37.6 ms | 39.6 ms |
| 65K | 141 ms | 149 ms |
A process starts in about 6 s; the first request of each new shape then costs one plan build (above). Peak GPU memory of a single plan is 1.7 GB at ≤ 512 tokens and 6.2 GB at 65K.
Intended use
- Fast, machine-consumable decisions grounded in the supplied text:
- policy and eligibility checks;
- verifying a claim against a document;
- extraction-style lookups;
- routing and triage among a few options;
- evidence checks over logs, code or records.
- Use the probabilities: act automatically on confident answers, and send uncertain ones to a larger model or a human.
Out of scope:
- text generation or chat;
- answering from world knowledge without supporting text;
- tasks whose correctness depends on multi-step reasoning;
- unsupervised high-stakes decisions (medical, legal, financial, safety, employment) or any use that needs guaranteed correctness.
Limitations
- Long, crowded documents. Exact lookup among near-identical decoys falls from 1.00 at ≤ 512 tokens to about 0.45 at 65K. As an example, in a 38K-token log of 1,100 near-identical INFO lines with one ERROR line, the model named the error's cause correctly ("disk full", 0.82) but answered "did any job fail?" with no (p = 0.29) and named the wrong worker. Choice among 255 options falls to about 0.30 beyond 32K.
- Calibration under distribution shift. Calibration error is about 0.07–0.08 on data unlike the training data, against 0.01–0.03 in distribution, so the model is somewhat overconfident on unfamiliar inputs. Calibration is weakest for Choice questions over long documents (e.g. multiple-choice questions about long stories).
- Option order. 9.4% of Choice decisions change when the options are permuted.
- No abstention. The model always returns a distribution, even when the document holds no evidence for the question. For example, an empty document gave "yes" with p = 0.72. Check that the supplied text actually covers what you ask.
- Limits. At most 65,536 total tokens, 8 questions, 255 options, 26 levels, and 8,192 tokens per question including its options. Digits are tokenized one by one, so number-heavy logs fit about 95K characters.
- Hardware. The FP8 runtime needs Blackwell (sm_120). The latency floor is about 1.7 ms per request (2.7 ms at 512 tokens).
- Bias and safety. The model inherits the base model's biases. Its decisions reflect patterns in the supplied text; they are not authoritative judgements.
Licence and attribution
Apache-2.0 (LICENSE). Derived from Qwen3.5-0.8B-Base by the Qwen team (Apache-2.0); see NOTICE.
Citation
@misc{qsiv2026,
title = {QSIV: a non-autoregressive typed-decision model},
author = {{QSIV authors}},
year = {2026},
note = {Open weights, derived from Qwen3.5-0.8B-Base}
}
- Downloads last month
- 161
Model tree for djinn-san/QSIV
Base model
Qwen/Qwen3.5-0.8B-Base









