QSIV

QSIV thumbnail

QSIV is a non-autoregressive decision model. You give it a document (code, logs, a policy, a contract, JSON) of up to 65,536 tokens and 1–8 typed questions about it. In a single forward pass it returns a calibrated probability distribution for each question. It does not generate text.

Question type You supply You get
Noul (yes/no) a question P(yes)
Choice a question and 2–255 options a probability for each option
Score a question and 2–26 ordered levels a probability for each level, plus the expected level

Questions are independent. Adding, removing or reordering the questions in a request never changes another question's answer (bitwise identical), so one request can carry several unrelated checks.

  • Developed by: the QSIV authors.
  • Model type: decision head on a hybrid Gated DeltaNet / attention transformer. 752,393,024 parameters, derived from Qwen3.5-0.8B-Base.
  • Weights: FP8 for production (qsiv_fp8.safetensors) and bf16 (model.safetensors), safetensors only.
  • Licence: Apache-2.0.

QSIV: up to 4,389× cheaper than LLMs and 21× cheaper than Jev per decision, $2 per billion input tokens, 1.76 ms per decision

QSIV vs. Doom: the model was faster than the game

We pointed QSIV at a live ViZDoom marine, with no eyes and no text generation, and asked for a fresh decision on every one of Doom's 35 ticks per second.

Then something unexpected happened: the game became the bottleneck.

Rendering a single Doom tic takes the engine about 3.1 ms at 640×480. QSIV decides in about 2.2 ms with a 4-question request and about 3.0 ms with 8 questions. The model was keeping pace with the game engine itself, and on lighter requests it was ahead.

So we had to speed the game up to keep the model fed. We moved the engine onto worker threads and dropped the render to 320×240 (about 1.0 ms per tic). Only then did the bottleneck swing back to QSIV, which, with 4–8 worker threads (each with its own CUDA stream over the same weights), held 16 live games at 35 Hz at once at roughly 650–710 decisions per second.

QSIV playing Doom: the right-hand panel shows the model's input, output and per-decision latency

Click the clip for the full video (62 s). The right-hand panel shows everything the model sees and says: the exact input document, every probability, the keys pressed, and the latency of every single decision on a live graph.

Numbers

Measured on one RTX PRO 6000 Blackwell, with the production FP8 runtime and CUDA-graph replay, one request at a time (batch size 1), 0 output tokens. Latency is the full decide() call (tokenization, forward pass and calibration) over 153,314 live decisions.

Decision latency p50 / p90 / p99 / max 2.45 / 3.18 / 3.34 / 6.5 ms
4-question request about 2.2 ms
8-question request about 3.0 ms
Fresh-document microbenchmark (p50 / p99) 2.26 / 3.03 ms
Decision rate about 400 decisions/s (the game needs 35/s)
Share of each 28.6 ms game tic about 10%
Game engine alone, 640×480 about 3.1 ms per tic
Game engine alone, 320×240 about 1.0 ms per tic
One model, many players 16 live games at 35 Hz, about 650–710 decisions/s
Peak GPU memory 4.3 GB

The whole pipeline (decision plus engine step) runs about 5× faster than real time.

At a glance

Measured on one RTX PRO 6000 Blackwell against two other open typed-decision models, Kev-0.8B (the same base model and size) and Laya Typed-Decisions. Every request carries a new document. Details and all numbers: benchmarks/COMPARISON.md.

Latency per request: QSIV 3–6× faster than Kev-0.8B

Speed-up over Kev-0.8B at every size, up to 9.5×

Throughput with 64 concurrent clients

Accuracy on long documents

Against Jev (TypeSafe's first System One model, available only as a cloud API). We did not measure Jev. Every Jev latency below was reported by others: by TypeSafe itself, by API routers (OpenRouter, OrcaRouter), by Every, and by the Kev project, which published its own Jev runs. Those Jev numbers therefore include the network. The QSIV numbers are our own measurements: svc.decide() inside the application, on a new document given as text (tokenization included), so there is no network at all. On requests of the Kev project's published Jev suites' size (~400 tokens), Jev's median is 222–232 ms and QSIV's is 3.0 ms: 74× faster. The fastest QSIV decision is 1.76 ms. Sources and conditions: benchmarks/JEV_LATENCY.md and benchmarks/jev_latency_sources.json.

A third party measured Jev at five document lengths (1K–32K tokens; jev-1.13 via OpenRouter, warm connection, huggingface.co/datasets/lisonallen/jev-ai-benchmark). At the same lengths QSIV is 10–77× faster: 4.9 ms against Jev's 375 ms at 1K tokens, and 72 ms against 734 ms at 32K.

At every document length: QSIV 10–77× faster than Jev

Requests of the same size: QSIV vs Jev

Self-hosting cost: ~$2 per billion input tokens ($0.002 per million), output free. QSIV is an open-weights model and 100% free to use under the Apache-2.0 licence—there are no API fees or subscription charges whatsoever. To provide an apples-to-apples comparison against commercial APIs like Jev ($42 per billion) and LLMs ($200–10,000 per billion), we measure what QSIV costs to self-host on rented hardware: on one rented RTX PRO 6000 Blackwell ($2.09 per hour on RunPod), the hardware operating cost works out to approximately $2 per billion input tokens across workflows. (Actual self-hosting costs vary across platforms and hardware setups—for example, Vast.ai median $1.51 per hour lowers this by 28%, and owned hardware costs even less). At this self-hosted rate, QSIV is 21× cheaper than Jev on every request, and 100× to 5,000× cheaper than the LLMs' published input prices, before counting the LLMs' output tokens.

A billion input tokens: $2 with QSIV, $42 with Jev, up to $10,000 with an LLM

per million input tokens per billion input tokens QSIV is cheaper by
QSIV (self-hosted) $0.002 $2
Jev (published) $0.042 $42 21×
GPT-5.6 Luna (published) $0.20 $200 100×
Claude Haiku 4.5 (published) $1.00 $1,000 500×
GPT-5.6 Terra (published) $2.00 $2,000 1,000×
Claude Fable 5.1 (implied by TypeSafe's 238× claim) $10.00 $10,000 5,000×

Per decision: 178–4,389× cheaper than LLMs. TypeSafe compares Jev with LLMs by the cost of each decision on its workflows, where an LLM also pays for output and reasoning (that is where its "444.6× cheaper" comes from). On the same published workflows, with QSIV's self-hosting cost applied to the same input, a million decisions cost $19 with QSIV, $390 with Jev and $3,310–120,000 with the LLMs. By TypeSafe's own 444.6× for Jev, QSIV is over 9,000× cheaper than LLMs. Jev's and the LLMs' costs are published by others, not measured by us. Method, every number and sources: benchmarks/PRICING.md.

A million decisions: $19 with QSIV, $390 with Jev, up to $120,000 with an LLM

What each open model costs to serve on the same GPU (QSIV, Kev-0.8B and Laya, all as HTTP servers):

Same GPU, same requests: QSIV costs 3× less to serve than Kev-0.8B

Quick start

Requirements: an NVIDIA Blackwell GPU (sm_120, e.g. RTX PRO 6000) with CUDA 13.0, and Python ≥ 3.12. The production runtime is FP8 Triton kernels replayed as CUDA graphs; other GPU generations are not supported by the FP8 path.

Coming soon: support for laptop GPUs, other desktop GPUs, Macs, and phone GPUs. An open-source release of the implementation is also coming soon.

pip install -U huggingface_hub        # provides the `hf` command
hf download djinn-san/QSIV --local-dir qsiv && cd qsiv
pip install -e .                       # torch 2.11, triton 3.6, transformers 5.14, safetensors, numpy
python examples/quickstart.py
from qsiv.runtime.service import QSIVService

svc = QSIVService("qsiv_fp8.safetensors", "calibration.json", tokenizer_dir="tokenizer")
svc.decide({
    "state": open("service.log").read(),
    "questions": [
        {"type": "noul",   "text": "Did the deploy at 14:02 fail?"},
        {"type": "choice", "text": "Which service failed first?", "options": ["api", "db", "cache"]},
        {"type": "score",  "text": "How severe is the incident?", "levels": ["none", "low", "medium", "high"]},
    ],
})
# [{'type': 'noul', 'p_true': ..., 'decision': ..., 'confidence': ...},
#  {'type': 'choice', 'choice': k, 'option': ..., 'probs': [...], 'confidence': ...},
#  {'type': 'score', 'level': k, 'level_label': ..., 'score': ..., 'uncertainty': ..., 'probs': [...], 'confidence': ...}]
  • First run and warm-up (measured on a fresh machine):
    • the very first python examples/quickstart.py takes about 25 s, because the Triton kernels compile; Triton caches them on disk, and later processes start in about 7 s (model load about 5 s);
    • inside a process, the first request of each new shape (state bucket × number of questions × question length) takes about 0.05–0.6 s to build and capture its CUDA-graph plan; every later request of that shape runs at the speeds below. Plans stay cached within a GPU-memory budget (QSIVService(..., plan_memory_gb=...), default min(24 GB, 40% of the GPU)): a mixed workload of 12 shapes from 64 to 65K tokens holds about 16 GB;
    • a server should pre-build the shapes it expects at startup, e.g. svc.warmup([(512, 1, 32), (512, 1, 64), (4096, 1, 64), (16384, 1, 64), (65536, 1, 64), (512, 8, 64)]), where each tuple is (state bucket, questions, question-block bucket), with buckets listed in config.json (a short yes/no question fits the 32-token bucket; a Choice question with its options usually needs 64 or more).
  • Integrity: verify the downloads with sha256sum -c qsiv_fp8.safetensors.sha256 model.safetensors.sha256.

Fast paths (all exact: same token ids, same outputs):

  • repeated question texts and recently seen long documents are tokenized once and cached;
  • longer requests are tokenized in parallel: the document (split only at positions where the tokenizer itself splits) and the questions go to the tokenizer's threads in one call;
  • a document you will query repeatedly can be tokenized once with ids = svc.encode_state(text) and then passed as svc.decide({"state_ids": ids, "questions": [...]}).

bf16 reference path. model.safetensors runs through the PyTorch model definition with qsiv.reference.ReferenceService (same interface; needs pip install -e ".[reference]"). It is meant for research, fine-tuning and porting. It is slower than the FP8 runtime, and its decisions match the FP8 runtime's on about 96% of 422 real held-out questions (mean largest probability difference 0.022; the differences come from FP8 rounding, concentrated on low-margin questions).

Files

File Contents
qsiv_fp8.safetensors production weights: FP8 E4M3 matrices with per-tensor scales, static activation scales in the header
model.safetensors the same model in bf16 (GDN gate parameters in fp32)
calibration.json temperatures per question type and per type × request-length band, applied by QSIVService
config.json, readout.json architecture, FP8 contract, request limits, plan buckets; question templates and answer labels
tokenizer/ Qwen3.5 tokenizer
qsiv/ inference package: FP8 Triton runtime, request rendering, model definition
examples/quickstart.py the example above
benchmarks/ BENCHMARKS.md (all results), COMPARISON.md (vs Kev-0.8B and Laya), JEV_LATENCY.md (vs Jev), PRICING.md (cost per token), raw result JSON

Benchmarks

Speed and document-reading results next to Kev-0.8B and Laya Typed-Decisions: benchmarks/COMPARISON.md.

All numbers are measured on these weights through the production FP8 runtime with the shipped calibration. The baseline is the untouched base model, run through the same runtime on the same samples. Full tables: benchmarks/BENCHMARKS.md.

Held-out benchmark evaluation: QSIV beats the base model on all 14 evaluation sources (gains from +0.003 to +0.548). Mean accuracy is 0.82, against 0.48 for the base model.

Real code and logs, never seen in training:

task base model QSIV
parameter binding / direct callers in real Python packages 0.56 / 0.11 0.96 / 0.86
source IP / count of failed SSH logins (Loghub 2.0) 0.75 / 0.10 0.95 / 0.41

Long context (synthetic probes; accuracy at 1K / 16K / 65K total tokens):

task 1K 16K 65K
retrieve a fact from the start of the document 1.00 1.00 1.00
apply a later update that overrides an earlier value 0.98 0.65 0.85
exact key lookup among near-identical decoys 0.98 0.53 0.45
Choice among many options (26 → 255) 1.00 0.72 0.30

Calibration (expected calibration error, lower is better): 0.010 on independent in-distribution test data and 0.081 on shifted data. By type: yes/no 0.013, Score 0.026, Choice 0.087.

Latency (RTX PRO 6000 Blackwell). Every request carries a new document and new questions, and the time is the complete svc.decide() call: tokenization, model and post-processing (median, cold L2 cache):

request tokens 1 question 8 questions
~60 1.76 ms 2.6 ms
~500 3.0 ms 3.9 ms
4K 10.8 ms 12.6 ms
16K 37.6 ms 39.6 ms
65K 141 ms 149 ms

A process starts in about 6 s; the first request of each new shape then costs one plan build (above). Peak GPU memory of a single plan is 1.7 GB at ≤ 512 tokens and 6.2 GB at 65K.

Intended use

  • Fast, machine-consumable decisions grounded in the supplied text:
    • policy and eligibility checks;
    • verifying a claim against a document;
    • extraction-style lookups;
    • routing and triage among a few options;
    • evidence checks over logs, code or records.
  • Use the probabilities: act automatically on confident answers, and send uncertain ones to a larger model or a human.

Out of scope:

  • text generation or chat;
  • answering from world knowledge without supporting text;
  • tasks whose correctness depends on multi-step reasoning;
  • unsupervised high-stakes decisions (medical, legal, financial, safety, employment) or any use that needs guaranteed correctness.

Limitations

  • Long, crowded documents. Exact lookup among near-identical decoys falls from 1.00 at ≤ 512 tokens to about 0.45 at 65K. As an example, in a 38K-token log of 1,100 near-identical INFO lines with one ERROR line, the model named the error's cause correctly ("disk full", 0.82) but answered "did any job fail?" with no (p = 0.29) and named the wrong worker. Choice among 255 options falls to about 0.30 beyond 32K.
  • Calibration under distribution shift. Calibration error is about 0.07–0.08 on data unlike the training data, against 0.01–0.03 in distribution, so the model is somewhat overconfident on unfamiliar inputs. Calibration is weakest for Choice questions over long documents (e.g. multiple-choice questions about long stories).
  • Option order. 9.4% of Choice decisions change when the options are permuted.
  • No abstention. The model always returns a distribution, even when the document holds no evidence for the question. For example, an empty document gave "yes" with p = 0.72. Check that the supplied text actually covers what you ask.
  • Limits. At most 65,536 total tokens, 8 questions, 255 options, 26 levels, and 8,192 tokens per question including its options. Digits are tokenized one by one, so number-heavy logs fit about 95K characters.
  • Hardware. The FP8 runtime needs Blackwell (sm_120). The latency floor is about 1.7 ms per request (2.7 ms at 512 tokens).
  • Bias and safety. The model inherits the base model's biases. Its decisions reflect patterns in the supplied text; they are not authoritative judgements.

Licence and attribution

Apache-2.0 (LICENSE). Derived from Qwen3.5-0.8B-Base by the Qwen team (Apache-2.0); see NOTICE.

Citation

@misc{qsiv2026,
  title  = {QSIV: a non-autoregressive typed-decision model},
  author = {{QSIV authors}},
  year   = {2026},
  note   = {Open weights, derived from Qwen3.5-0.8B-Base}
}
Downloads last month
161
Safetensors
Model size
0.8B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for djinn-san/QSIV

Finetuned
(135)
this model