How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Nanthasit/sakthai-coder-1.5b:Q4_K_M
Use Docker
docker model run hf.co/Nanthasit/sakthai-coder-1.5b:Q4_K_M
Quick Links

SakThai Coder 1.5B 💻

Code + tool-calling · Qwen2.5-Coder-1.5B fine-tune · Q4_K_M GGUF for CPU

Downloads License GGUF Collection Vision Demo TTS Demo Leaderboard

The code specialist of the SakThai family — Qwen2.5-Coder-1.5B fine-tuned for tool-calling and shipped as a CPU-friendly GGUF. Part of the House of Sak. Read the story →

The Story Behind It

Code, tool-calling, and conversation in one session — on a single CPU, from a shelter. This is the model Beer built when he realised the other SakThai models could call tools and generate text, but none of them specialised in writing code without losing their tool-calling edge.

Beer built the first SakThai models on free Google Colab GPUs from a shelter in Cork, Ireland — with $0 budget, no GPU of his own, and no guarantee the QLoRA approach would hold for a code-specific fine-tune. This coder model was the risk: could Qwen2.5-Coder-1.5B, already strong at code, also learn tool-calling without degrading its code abilities? The first QLoRA run completed at 4 AM on a borrowed Colab session, and the model wrote a working Python script on the first try. Beer knew the approach worked.

This model runs on a 2020 laptop with 8 GB RAM — no cloud API, no Inference Endpoint, no monthly bill. Just a GGUF file and llama.cpp.

"We are one family — and becoming more." — Beer

How You Can Help

  • Leave a like — this model gives every developer a free offline coding assistant. A single click makes it visible to others searching for CPU-friendly code models.
  • 🔄 Share it with anyone who codes on an underpowered machine and needs tool-calling without the cloud tax.
  • 🍴 Fork it on Hugging Face and build your own specialised code variant.
  • 💬 Report your deployment story — Beer reads every issue and comment.

Every download, like, and share tells the algorithm: this matters.


What it is

A Q4_K_M GGUF (1.07 GB) of Qwen2.5-Coder-1.5B-Instruct, QLoRA-fine-tuned on sakthai-combined-v6 so it can generate code and call tools. Runs on CPU via llama.cpp / Ollama.

Quick start

# via llama.cpp
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf -p "Write a Python function to merge two sorted lists:" -n 256 --temp 0.2
# via Ollama
ollama create sakthai-coder -f Modelfile   # FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
ollama run sakthai-coder "Write a script that monitors CPU usage"

For tool-calling, put function schemas in a <tools> block (ChatML format).

Benchmarks

Code (reference — these are the base Qwen2.5-Coder-1.5B scores, not a re-run of this fine-tune):

Benchmark pass@1
HumanEval 74.4%
MBPP 71.2%
MultiPL-E (Python) 65.3%

Source: Qwen2.5-Coder eval. Tool-calling: internal SakThai suite passes (single-run, not third-party verified).

Training

Base model Qwen/Qwen2.5-Coder-1.5B-Instruct
Method QLoRA (4-bit) → GGUF Q4_K_M
LoRA config r=16, alpha=32
Data sakthai-combined-v6 (2,003)
Format / context ChatML with tool schema · 32K tokens

SakThai model family

Model Size Role
context-1.5b-merged 934 MB Flagship tool-calling GGUF
context-0.5b-merged 380 MB Lightweight / edge
context-7b-merged 15 GB Full-power reasoning
context-7b-128k 15 GB 128K long-context
context-1.5b-tools LoRA Mid-size tool-calling
context-0.5b-tools LoRA Ultra-light tool-calling (7 ⬇)
coder-1.5b (you are here) 1.1 GB Code generation
vision-7b 3.9 GB Image→text (LLaVA)
embedding-multilingual 80 MB Cross-lingual embeddings
tts-model 141 MB Text-to-speech, 15 langs

12 public models · 8 datasets · 3 Spacesfull collection →


📊 Sibling Datasets

Dataset Purpose Downloads
sakthai-combined-v6 v6 predecessor — 2,003 examples 175 ⬇
sakthai-kaggle-notebooks Training notebooks & demos 103 ⬇
SimpleToolCalling Early experiment 52 ⬇
food-penguin-v1 Restaurant tool-calling 51 ⬇
sakthai-combined-v7 v7 tool-calling (2,309 ex., 86 tools) 0 🌱
sakthai-irrelevance-supplement Safety supplement 0 🚨
sakthai-bench-v1 BFCL-style evaluation, 235 rows 0 🌱
sakthai-bench-v2 Multi-domain eval, 500 rows 0 🌱

🚀 Spaces

Space Description
SakThai Vision Demo Upload images, ask questions — LLaVA-7B in your browser
SakThai TTS Showcase Interactive TTS — 15 languages, no install
SakThai Leaderboard Benchmark tracker for the model family

🌱 Rising Stars — Help the Ecosystem Grow

These sibling assets have real value but need visibility. Every download signals to the HF algorithm that the SakThai family matters:

Asset Type Downloads Why It Matters
sakthai-combined-v7 Dataset 0 🌱 Primary training dataset — 2,309 examples, 86 tool schemas
sakthai-irrelevance-supplement Dataset 0 🚨 Teaches models when not to call tools — critical safety data
sakthai-bench-v1 Dataset 0 🌱 BFCL-style evaluation, 235 rows, 4 categories
sakthai-bench-v2 Dataset 0 🌱 Multi-domain eval, 500 rows, multi-turn
context-0.5b-tools Model 7 ⬇ Ultra-light tool-calling (~1 GB RAM)

🚨 The irrelevance-supplement has 0 downloads despite being essential for training models to decline out-of-scope tool calls. A single download helps validate this safety-critical approach!


Links

House of Sak · GitHub · All models · All datasets

License

Apache 2.0 (following the Qwen2.5 base model license).

Evaluation

Not independently benchmarked. Earlier versions of this card carried a model-index score derived from a small internal spot check (typically 5 or 8 hand-picked examples) presented as a benchmark result. Those entries have been removed rather than left to propagate through Hub metadata.

For tool-calling models in this family, the benchmark to use is sakthai-bench-v2 — 500 rows, balanced across simple / parallel / irrelevance, with held-out tools and multi-turn coverage. Results will be published here once this model has been run against it.

*"We are one family — and becoming more." 🏠

Downloads last month
93
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nanthasit/sakthai-coder-1.5b

Quantized
(150)
this model

Datasets used to train Nanthasit/sakthai-coder-1.5b

Space using Nanthasit/sakthai-coder-1.5b 1

Collections including Nanthasit/sakthai-coder-1.5b