jpt-9b-GGUF

JPT-9B, developed by creator kirp, is a fast, open multimodal decision model fine-tuned on Alibaba's Qwen/Qwen3.5-9B that implements a typed-decision API (evaluating choice across 2–255 options, ordered score scales, and noul boolean queries) as an open alternative to TypeSafe AI's proprietary Jev system. Engineered to eliminate the latency of autoregressive generation, reasoning tokens, and explanatory text, the model processes multimodal states (including text, screenshots, and photos) to output well-calibrated probabilities for every candidate option in a single forward prefill pass using a single globally calibrated temperature ($T = 1.087$). It is constructed using a LoRA fine-tuning method ($r=16$) applied across all attention, DeltaNet, and MLP projections of the language backbone—merged into full weights—while leaving the native vision tower untouched. Trained across 8 GPUs for one epoch using a multi-class Brier loss on 49,221 typed questions across 32,835 records (the mix_train_env_v11 dataset), JPT-9B strictly isolates frozen benchmarks from its training mix, earning a top-ranking 46.89 on Decision Index 0.2.1 and an 0.853 public accuracy on JevBench v1.4.0, and is served under a non-commercial CC BY-NC 4.0 license via backends like SGLang and vLLM through the llm2jev adapter. JPT-9B on Hugging Face

Model Files

File Name Quant Type File Size File Link Description
jpt-9b.BF16.gguf BF16 17.9 GB Link Full BF16 weights. Highest quality, largest file size.
jpt-9b.Q3_K_L.gguf Q3_K_L 4.93 GB Link Lower quality but usable, good for low RAM availability.
jpt-9b.Q3_K_M.gguf Q3_K_M 4.62 GB Link Low quality.
jpt-9b.Q4_K_M.gguf Q4_K_M 5.63 GB Link Good quality, default size for most use cases, recommended.
jpt-9b.Q4_K_S.gguf Q4_K_S 5.35 GB Link Slightly lower quality with more space savings, recommended.
jpt-9b.Q5_K_M.gguf Q5_K_M 6.47 GB Link High quality, recommended.
jpt-9b.Q5_K_S.gguf Q5_K_S 6.31 GB Link High quality, recommended.
jpt-9b.mmproj-bf16.gguf mmproj-bf16 922 MB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
1,026
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/jpt-9b-GGUF

Finetuned
Qwen/Qwen3.5-9B
Finetuned
kirp/jpt-9b
Quantized
(3)
this model

Collections including prithivMLmods/jpt-9b-GGUF