Instructions to use KikoCis/MiniCPM5-2B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use KikoCis/MiniCPM5-2B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0 # Run inference directly in the terminal: llama cli -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Use Docker
docker model run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0
- LM Studio
- Jan
- vLLM
How to use KikoCis/MiniCPM5-2B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "KikoCis/MiniCPM5-2B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "KikoCis/MiniCPM5-2B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0
- Ollama
How to use KikoCis/MiniCPM5-2B-GGUF with Ollama:
ollama run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0
- Unsloth Desktop
- Pi
How to use KikoCis/MiniCPM5-2B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "KikoCis/MiniCPM5-2B-GGUF:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use KikoCis/MiniCPM5-2B-GGUF with Docker Model Runner:
docker model run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0
- Lemonade
How to use KikoCis/MiniCPM5-2B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull KikoCis/MiniCPM5-2B-GGUF:Q8_0
Run and chat with the model
lemonade run user.MiniCPM5-2B-GGUF-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use KikoCis/MiniCPM5-2B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default KikoCis/MiniCPM5-2B-GGUF:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use KikoCis/MiniCPM5-2B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf KikoCis/MiniCPM5-2B-GGUF:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "KikoCis/MiniCPM5-2B-GGUF:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
BF16 โโโโโโโโโโโโโโโ token_embd โโ bf16 โ KLD 0.000616 blk.0..41 โโ q8_0 โโโโถ top-1 98.67 % output โโ bf16 โ PPL +0.0109 โโโโโโโโโโโโโโโโโโโโ
FORMAT GGUF |
SIZE 3.18 / 2.68 GB |
ARCH llama, 42 layers |
CONTEXT 131,072 |
QUANTS Q8_0-XL ยท Q8_0 |
VALIDATION KLD / PPL / top-1 |
RUNS ON llama.cpp ยท Ollama |
LICENSE Apache-2.0 |
MiniCPM5-2B โ Q8_0-XL + Q8_0 GGUF
8-bit GGUFs of OpenBMB's MiniCPM5-2B, each one measured against the BF16 original. Q8_0-XL keeps the token embeddings and the output head in BF16 and puts every transformer block in Q8_0: 11 % lower KLD and 26 % less perplexity drift than a plain Q8_0, for +0.5 GB. Runs in ~3โ4 GB of RAM at 8K context.
This is OpenBMB's model, our quantization. OpenBMB also publishes their own GGUFs; we measured their Q8_0 alongside ours (table below) and it is equivalent to our plain Q8_0. What this repo adds is the XL variant and the numbers.
โ Recommended files
| Use case | File | Notes |
|---|---|---|
| Best fidelity at 8 bits | MiniCPM5-2B-Q8_0-XL.gguf |
Lowest KLD and PPL drift here. Recommended. |
| Smallest 8-bit | MiniCPM5-2B-Q8_0.gguf |
Standard Q8_0, 0.5 GB lighter. |
๐ฆ Files
| Quant | Bits (blocks / embed + head) | File size |
|---|---|---|
| Q8_0-XL | 8 / 16 | 3.18 GB |
| Q8_0 | 8 / 8 | 2.68 GB |
The embedding table and output head are unusually large for a 2B model (130,560-token vocabulary ร 2,048 = 1 GB in BF16 between the two), which is why keeping them in BF16 costs 0.5 GB and is where most of the remaining 8-bit error lives.
๐ Metrics โ fidelity vs the BF16 reference
KLD (KullbackโLeibler divergence, nats): how far the quant's next-token distribution drifts from BF16. Top-1 = how often the quant's most likely token is the same as BF16's. PPL ฮ = perplexity above BF16 (11.196).
| File | Size GB | PPL | PPL ฮ | KLD mean | KLD median | KLD p95 | KLD p99 | Top-1 match |
|---|---|---|---|---|---|---|---|---|
| BF16 reference | 5.04 | 11.196 | 0 | 0 | 0 | 0 | 0 | 100 % |
| Q8_0-XL | 3.18 | 11.207 | +0.0109 | 0.000616 | 0.000375 | 0.00193 | 0.00444 | 98.67 % |
| Q8_0 | 2.68 | 11.211 | +0.0147 | 0.000692 | 0.000451 | 0.00209 | 0.00463 | 98.57 % |
| Q8_0 โ OpenBMB's official file, for reference | 2.68 | 11.211 | +0.0147 | 0.000692 | 0.000450 | 0.00209 | 0.00464 | 98.56 % |
Standard errors: ยฑ0.000005 on mean KLD, ยฑ0.003 on PPL ฮ, ยฑ0.05 points on top-1. The XL gap is well outside them.
Machine-readable: metrics/quant-summary-with-kld.json / .csv. Per-file detail in reports/.
๐ Chart
๐งฎ Will it fit?
Memory โ file size + KV cache. MiniCPM5-2B uses grouped-query attention (2 KV heads), so long context is cheap:
| Context | KV cache (F16) | Q8_0-XL total | Q8_0 total |
|---|---|---|---|
| 8K | 0.35 GB | ~3.6 GB | ~3.1 GB |
| 32K | 1.4 GB | ~4.6 GB | ~4.1 GB |
| 128K (native) | 5.6 GB | ~8.9 GB | ~8.4 GB |
๐ How to run it
# llama.cpp โ the GGUF carries OpenBMB's chat template (tools + reasoning)
llama-server -m MiniCPM5-2B-Q8_0-XL.gguf -c 32768 --jinja --temp 1.0 --top-p 0.95
# Ollama โ Modelfiles for 8K / 32K / 128K are included
ollama run hf.co/KikoCis/MiniCPM5-2B-GGUF:Q8_0-XL
ollama create minicpm5 -f Modelfile-32k
Sampling: temperature 1.0, top_p 0.95 โ OpenBMB's own defaults (generation_config.json).
Reasoning: the model thinks before answering; Ollama returns that in a separate thinking field, llama-server in reasoning_content with --jinja.
Tools: native function calling through the embedded chat template โ use the tools parameter of the OpenAI-compatible API.
โ ๏ธ Good to know
- At 8 bits both files are already very close to BF16; the XL gain is real but small in absolute terms. Pick XL if you have 0.5 GB to spare, plain Q8_0 otherwise.
- No weights were changed: this is a faithful quantization of OpenBMB's release.
- We have not run agentic or task benchmarks on these files โ the numbers above measure fidelity to the original model, not its capability. For capability, see OpenBMB's card.
๐ Evaluation methodology
- What: KLD, perplexity and top-1 agreement of each GGUF against the BF16 GGUF converted from the same weights.
- Data: wikitext-2 test split (
wiki.test.raw), first 64 chunks of 2,048 tokens. - Tool:
llama-perplexity --kl-divergence-base(BF16 logits) then--kl-divergenceper file. - Source:
openbmb/MiniCPM5-2Bat revisionf97400052a43d642bbc6e9975e2397e3ae6a6b52. - Load test: both files loaded and answered correctly in llama.cpp and Ollama (chat template, reasoning split).
- Date: 2026-10-10.
๐ Provenance & reproducibility
- Exact commands:
scripts/build.shโ download at the pinned revision โ BF16 + Q8_0 convert โ Q8_0-XL quantize โ KLD/PPL. - Checksums:
reports/artifact-sha256sums.txtโshasum -a 256 -cto verify. - Per-file results:
reports/,metrics/.
๐๏ธ Changelog
- 2026-10-10 โ v1: Q8_0-XL and Q8_0, with KLD/PPL/top-1 vs BF16 and OpenBMB's Q8_0 as a reference row.
๐ Credit & license
Model, weights and training: ยฉ OpenBMB โ openbmb/MiniCPM5-2B. Quantization, measurements and Modelfiles: KikoCis. Apache-2.0, same as upstream.
- Downloads last month
- -
8-bit
Model tree for KikoCis/MiniCPM5-2B-GGUF
Base model
openbmb/MiniCPM5-2B
