How to use from the
Use from the
llama-cpp-python library
# !pip install llama-cpp-python

from llama_cpp import Llama

llm = Llama.from_pretrained(
	repo_id="Nanthasit/sakthai-coder-1.5b",
	filename="qwen2.5-coder-1.5b-instruct-q4_k_m.gguf",
)
llm.create_chat_completion(
	messages = [
		{
			"role": "user",
			"content": "What is the capital of France?"
		}
	]
)

SakThai Coder 1.5B 💻

Code + tool-calling · Qwen2.5-Coder-1.5B fine-tune · Q4_K_M GGUF for CPU

Downloads License GGUF Collection

The code specialist of the SakThai family — Qwen2.5-Coder-1.5B fine-tuned for tool-calling and shipped as a CPU-friendly GGUF. Part of the House of Sak. Read the story →

What it is

A Q4_K_M GGUF (1.07 GB) of Qwen2.5-Coder-1.5B-Instruct, QLoRA-fine-tuned on sakthai-combined-v6 so it can generate code and call tools. Runs on CPU via llama.cpp / Ollama.

Quick start

# via llama.cpp
wget https://huggingface.co/Nanthasit/sakthai-coder-1.5b/resolve/main/qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
./llama-cli -m qwen2.5-coder-1.5b-instruct-q4_k_m.gguf -p "Write a Python function to merge two sorted lists:" -n 256 --temp 0.2
# via Ollama
ollama create sakthai-coder -f Modelfile   # FROM ./qwen2.5-coder-1.5b-instruct-q4_k_m.gguf
ollama run sakthai-coder "Write a script that monitors CPU usage"

For tool-calling, put function schemas in a <tools> block (ChatML format).

Benchmarks

Code (reference — these are the base Qwen2.5-Coder-1.5B scores, not a re-run of this fine-tune):

Benchmark pass@1
HumanEval 74.4%
MBPP 71.2%
MultiPL-E (Python) 65.3%

Source: Qwen2.5-Coder eval. Tool-calling: internal SakThai suite passes (single-run, not third-party verified).

Training

Base model Qwen/Qwen2.5-Coder-1.5B-Instruct
Method QLoRA (4-bit) → GGUF Q4_K_M
LoRA config r=16, alpha=32
Data sakthai-combined-v6 (2,003)
Format / context ChatML with tool schema · 32K tokens

SakThai model family

Model Size Role
context-1.5b-merged 934 MB Flagship tool-calling GGUF
context-0.5b-merged 380 MB Lightweight / edge
context-7b-merged 15 GB Full-power reasoning
context-7b-128k 15 GB 128K long-context
context-{7b,1.5b,0.5b}-tools LoRA Tool-calling adapters
coder-1.5b 1.1 GB Code generation ⬅
vision-7b 3.9 GB Image→text (LLaVA)
embedding-multilingual 80 MB Cross-lingual embeddings
tts-model 141 MB Text-to-speech, 15 langs

12 models · 5 datasets · 3 Spacesfull collection →

Links

House of Sak · GitHub · All models

License

Apache 2.0 (following the Qwen2.5 base model license).

Evaluation

Not independently benchmarked. Earlier versions of this card carried a model-index score derived from a small internal spot check (typically 5 or 8 hand-picked examples) presented as a benchmark result. Those entries have been removed rather than left to propagate through Hub metadata.

For tool-calling models in this family, the benchmark to use is sakthai-bench-v2 — 500 rows, balanced across simple / parallel / irrelevance, with held-out tools and multi-turn coverage. Results will be published here once this model has been run against it.

Downloads last month
70
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nanthasit/sakthai-coder-1.5b

Quantized
(150)
this model

Datasets used to train Nanthasit/sakthai-coder-1.5b

Space using Nanthasit/sakthai-coder-1.5b 1

Collections including Nanthasit/sakthai-coder-1.5b