Instructions to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- jev-style
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with jev-style:
pip install jev-style # GGUF builds score through llama.cpp: build the jev-score binary once hf download chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF build_jev_score.sh jev_score.cpp --local-dir jev-score export JEV_SCORE_BIN=$(sh jev-score/build_jev_score.sh /path/to/llama.cpp | tail -n 1)
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Use Docker
docker model run hf.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with Ollama:
ollama run hf.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with Docker Model Runner:
docker model run hf.co/chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
- Lemonade
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Jev-Style-0.8B-Decision-v3-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Jev-Style-0.8B-Decision-v3-GGUF
Jev-Style decision series: v1 · 2B → v2 · 2B → v3 · 0.8B · GitHub: jev-style · Website: jevstyle.com · Collection: all v3 builds and demos
Run it locally, inside your agents: github.com/lawrence3699/jev-style serves this model behind a systemone-compatible API with a Playground, and adds six agent skills (
npx skills add lawrence3699/jev-style), a Claude Code guard hook and MCP tools.pip install "jev-style", thenjev-style serve --backend gguf --scorer /path/to/jev-score, or in Python:from jev_style import JevStyle, noul js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF", quant="Q4_K_M", scorer="/path/to/jev-score") js.decide("I was charged twice.", {"billing": noul("This is about billing.")})Build the
jev-scorescorer once withbuild_jev_score.shfrom this repo (see Quick start).
Jev-style decisions on your laptop. These are the GGUF builds of Jev-Style-0.8B-Decision-v3 for llama.cpp: 0.53 GB in 4-bit (Q4_K_M), with the same decision as full precision on 240 of 240 parity rows. Full results, protocols, training data and licences are on the main model card.
| Beyond its training data | Jev-Style v3 · 0.8B | Best official Laya |
|---|---|---|
| Banking77, 77 intents (never trained) ↑ | 68.2% | 49.2% |
| MASSIVE intent, 37 held-out languages ↑ | 65.5% | 36.1% |
| tweet_topic, zero-shot ↑ | 75.5% | 63.2%¹ |
| JevBench v1.4.1, 231 public items, zero-shot ↑ | 64.1% | 58.4%² |
| Runs on your own machine | Yes, 0.53 GB (4-bit) | Yes |
| Longest input per call | 25,600 tokens | 1,024 by default³ |
Laya: the best of its three official checkpoints, re-run by us on identical rows with their shipped temperatures; paired 95% CIs exclude zero for Banking77 and MASSIVE. ¹ English Laya, as published by the elcronos study. ² Laya's score as published on the JevBench board; it lies inside v3's 95% CI, so this lead is a point estimate. ³ Default input budget in the Laya README: 1,024 tokens for the multilingual and typed checkpoints, 512 for English. Jev (API) has higher accuracy than v3 on each of these sets where its accuracy is published. Protocols and confidence intervals: see the main model card.
Reads long documents in one call. Up to 25,600 tokens of input, 25× Laya's 1,024-token default and 25× our 2B v2's prompt. On 1,280 real 24K-token items v3 answers 98.3% correctly, and accuracy stays flat from 1K to 24K tokens (preregistered claim, passed).
Also: ahead of Laya multilingual in 51 of 51 languages · +2.6 points over Laya's typed checkpoint on typed decisions, trained on the same split (how to read that number).
Files
| File | Quantization | Size | Same decision as PyTorch FP32 (240 parity rows) | Prompts of about 16K / 25.6K tokens |
|---|---|---|---|---|
Jev-Style-0.8B-Decision-v3-F16.gguf |
F16 | 1.52 GB | 240 / 240 | 6 / 6 |
Jev-Style-0.8B-Decision-v3-Q8_0.gguf |
Q8_0 | 0.81 GB | 240 / 240 | 6 / 6 |
Jev-Style-0.8B-Decision-v3-Q4_K_M.gguf |
Q4_K_M | 0.53 GB | 240 / 240 | 6 / 6 |
Top-1 agreement with the PyTorch FP32 reference on a 240-row mixed parity fixture (training-pool rows, 22 categories, English and Chinese) plus 6 extra long prompts. These rows test agreement between formats, not accuracy. Sizes are the exported files (GB = 10^9 bytes). F16 is the runtime's default and the backend used for the latency figures.
Quick start
The files are standard Qwen3.5 text models, so any recent llama.cpp loads them. Chat or text generation does not give you the model's decisions, though. Decisions are read at one verdict slot per option, and the bundled scorer does exactly that.
pip install -U huggingface_hub
hf download chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF --local-dir jev-v3-gguf
cd jev-v3-gguf
pip install -r requirements.txt # tokenizers, numpy
# Build the scorer against llama.cpp (tested at commit 441df11f65ea0b6d0c72965aaf70c8241070ddcb or later).
git clone https://github.com/ggml-org/llama.cpp
git -C llama.cpp checkout 441df11f65ea0b6d0c72965aaf70c8241070ddcb
sh build_jev_score.sh llama.cpp # -> ./build/jev-score (Metal on macOS)
# Linux + CUDA: LLAMA_CMAKE_FLAGS="-DGGML_CUDA=ON" sh build_jev_score.sh llama.cpp
python jev_style_decision_gguf.py --quant Q4_K_M \
--state "The film was excellent." \
--question "What is the sentiment of this review?" \
--options '["negative", "positive"]' --category general_sentiment
From Python, the API is the same as the main repository's runtime:
from jev_style_decision_gguf import JevStyleDecisionGGUF
m = JevStyleDecisionGGUF(".", quant="F16") # or "Q8_0", "Q4_K_M"
r = m.decide(
{"ticket": "I was charged twice for my subscription this month.", "customer_tier": "pro"},
"Which team should handle this ticket?",
options={"billing": "payments, invoices, refunds", "technical": "bugs and outages", "sales": "new purchases"},
category="theme_routing",
)
print(r["answer"], r["probabilities"])
m.close()
jev-score(jev_score.cpp) is a small libllama program that runs as a JSON-lines process. It requests logits only at the slot positions. With a shared prefix it decodes the state once and scores several questions on copies of it.decide_manysends all questions about one state in one request. The default,many_mode="exact", returns exactly what onedecidecall per question returns; it shares the state in whole 1,024-token blocks, so it saves time from 1,024-token states on.many_mode="batched"reads the whole state once and scores all questions together (the setting of the latency chart); its probabilities differed fromdecideby at most 0.002 in our tests, and a near-tied top answer can change.- The runtime opens a 32,768-token context, which covers the 25,600-token input limit plus the question part. The whole input may be up to 25,600 tokens, and the question, options and readout up to 2,048. Over-budget inputs raise an error, and nothing is truncated.
- Long option lists (added 2026-09-26). When a choice question's options do not fit the 2,048-token
budget together, the runtime scores them in option chunks. Each chunk is an ordinary question with the same
text and a contiguous slice of the options, and the chunks are as few and as even as possible. The per-option
scores of all chunks then go through one softmax (
option_chunksin the result). Questions that fit are unchanged: on a 476-request test set, every one of them came back bit-identical to the previous runtime. Use--no-split-options(orsplit_options=False) to get the old error instead. - The input format, the readout and the calibration temperatures are described on the main card.
Results and speed
- 4-bit, 0.53 GB, same calls. The Q4_K_M file matches PyTorch FP32 on 240 of 240 parity rows, plus 6 of 6 prompts at about 16K and 25.6K tokens, and it is about 2.4× smaller than the 2B v2's Q4_K_M (1.27 GB).
- Up to 4.6× faster than a Laya-architecture engine when 10 questions share one 4K-token state (1,381 ms vs
6,364 ms with
many_mode="batched"; the engine is our round-1 MacLaya-4K, one call per question, not an official Laya checkpoint), because in that mode the bundled scorer reads the state once. It also answers questions about 8K-token states in 2.3 to 2.6 s. - 79.2% on 2,000 typed decisions, +2.6 points over Laya's typed checkpoint trained on the same split (paired 95% CI +1.0 to +4.2) and +5.7 over the 2B v2. In-domain, so it measures agreement with the dataset's teacher labels; see reading the typed number.
Latency: untrained identical-architecture Qwen3.5-0.8B export on llama.cpp GGUF F16, one call per state with all questions scored together (many_mode="batched"); comparison engine = round-1 MacLaya-4K, our own fine-tune of the Laya multilingual architecture (FP32 on Apple MPS, 4,096-token budget, one call per question), not an official Laya checkpoint; Apple M1 Max 64 GB, warm p50, idle run 2026-09-23. v3 parity rows are drawn from the training pool; 2B v1/v2 quantization numbers and sizes are as reported on their public GGUF cards (their own 500 held-out decisions), so no agreement gap is claimed. More protocol detail is on the main card.
Licence
Apache-2.0. Built on Qwen/Qwen3.5-0.8B (Apache-2.0). Some training data has restrictive or unclear terms, and some training rows are outputs of OpenAI and Anthropic models. See "Training data and licences" on the main card. Not affiliated with TypeSafe AI, Jev, the Laya authors or the Qwen team.
Contact
I welcome internship, employment, and research collaboration opportunities. Please contact me at yanchaoliang369@gmail.com.
欢迎提供实习、工作及科研合作机会,请邮件联系:yanchaoliang369@gmail.com。
- Downloads last month
- 2,090
4-bit
8-bit
16-bit
Model tree for chaoliangUNSW/Jev-Style-0.8B-Decision-v3-GGUF
Base model
Qwen/Qwen3.5-0.8B-Base

