Image-Text-to-Text
Transformers
GGUF
English
Thai
Chinese
qwen2-vl
vision-language
multimodal
image-understanding
tool-use
screenshot
vqa
grounding
sakthai
house-of-sak
cpu-inference
offline
Eval Results (legacy)
Eval Results
conversational
Instructions to use Nanthasit/sakthai-vision-7b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nanthasit/sakthai-vision-7b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Nanthasit/sakthai-vision-7b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Nanthasit/sakthai-vision-7b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Nanthasit/sakthai-vision-7b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Use Docker
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Nanthasit/sakthai-vision-7b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nanthasit/sakthai-vision-7b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- SGLang
How to use Nanthasit/sakthai-vision-7b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-vision-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-vision-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use Nanthasit/sakthai-vision-7b with Ollama:
ollama run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Nanthasit/sakthai-vision-7b with Docker Model Runner:
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- Lemonade
How to use Nanthasit/sakthai-vision-7b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Nanthasit/sakthai-vision-7b:Q4_K_M
Run and chat with the model
lemonade run user.sakthai-vision-7b-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -1,341 +1,188 @@
|
|
| 1 |
---
|
| 2 |
-
license:
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
pipeline_tag: image-to-text
|
| 6 |
library_name: transformers
|
|
|
|
| 7 |
tags:
|
| 8 |
-
- vision
|
| 9 |
- llava
|
|
|
|
| 10 |
- multimodal
|
| 11 |
-
- image-to-text
|
| 12 |
-
- visual-question-answering
|
| 13 |
- image-captioning
|
| 14 |
-
-
|
| 15 |
-
- llama-cpp
|
| 16 |
- sakthai
|
| 17 |
- house-of-sak
|
| 18 |
-
-
|
| 19 |
- cpu-inference
|
| 20 |
-
-
|
| 21 |
-
|
| 22 |
-
- privacy
|
| 23 |
-
- quantized
|
| 24 |
-
- transformers
|
| 25 |
-
- clip
|
| 26 |
-
- vicuna
|
| 27 |
-
- finetune
|
| 28 |
-
- safetensors
|
| 29 |
-
base_model: liuhaotian/LLaVA-1.5-7b
|
| 30 |
extra:
|
| 31 |
-
sibling: Nanthasit/sakthai-
|
| 32 |
-
formats: GGUF Q4_K_M
|
| 33 |
-
datasets:
|
| 34 |
-
- liuhaotian/LLaVA-Instruct-150K
|
| 35 |
-
- liuhaotian/LLaVA-Pretrain
|
| 36 |
---
|
| 37 |
|
| 38 |
-
|
| 39 |
-
|
| 40 |
<p align="center">
|
| 41 |
-
<
|
| 42 |
-
<
|
| 43 |
-
<img src="https://img.shields.io/badge/GGUF-Q4__K__M-orange" alt="GGUF"/>
|
| 44 |
-
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🏠-SakThai%20Family-6644cc" alt="Collection"/></a>
|
| 45 |
-
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🚀-Explore%20Family-47d147" alt="Collection"/></a>
|
| 46 |
-
<img src="https://img.shields.io/badge/benchmark-78.5%25%20VQAv2-success?logo=googlechrome" alt="VQAv2"/>
|
| 47 |
</p>
|
| 48 |
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
>
|
| 52 |
-
>
|
| 53 |
-
|
| 54 |
-
|
| 55 |
-
--
|
| 56 |
-
|
| 57 |
-
## Architecture
|
| 58 |
-
|
| 59 |
-
| Component | Detail |
|
| 60 |
-
|-----------|--------|
|
| 61 |
-
| **Model** | LLaVA-1.5-7B ([paper](https://arxiv.org/abs/2310.03744)) — Vicuna-7B-v1.5 LLM + CLIP ViT-L/14 vision encoder |
|
| 62 |
-
| **Vision encoder** | CLIP ViT-L/14, 336×336 input, 1024 embedding dim, 24 layers, 16 heads, ~427M params |
|
| 63 |
-
| **Language model** | LLaMA-based, ~7B parameters, 32 layers, 4096 hidden dim, 32 heads |
|
| 64 |
-
| **Projector** | 2-layer MLP (CLIP embedding → LLM token space), 4096 hidden dim |
|
| 65 |
-
| **Context window** | 4096 tokens |
|
| 66 |
-
| **Attention** | Causal self-attention (LLM) + bidirectional self-attention (vision encoder) |
|
| 67 |
-
| **Quantization** | GGUF Q4_K_M (4-bit group quantization, 256 group size) |
|
| 68 |
-
| **File size** | ~4 GB (LLM) + ~400 MB (mmproj) |
|
| 69 |
-
| **Original size** | ~6.7 GB (unquantized fp16) |
|
| 70 |
-
| **Chat template** | Vicuna-style with `<image>` placeholder |
|
| 71 |
-
| **License** | LLaMA 2 Community License |
|
| 72 |
-
|
| 73 |
-
LLaVA (Large Language and Vision Assistant) connects a pre-trained CLIP vision encoder to a Vicuna LLM via a lightweight projection MLP. The vision encoder processes images into embeddings, the projector maps them into the LLM's token space, and the LLM generates text conditioned on both the visual and textual input. This architecture enables multi-turn visual dialogue: once an image is encoded, you can ask follow-up questions without reprocessing the visual input.
|
| 74 |
-
|
| 75 |
-
This quantized GGUF build reduces the footprint from ~6.7 GB (fp16) to ~4 GB while retaining ~98–99% of the original accuracy.
|
| 76 |
-
|
| 77 |
-
> **Reference:** Liu et al., *"Visual Instruction Tuning"* (NeurIPS 2023). [arXiv:2310.03744](https://arxiv.org/abs/2310.03744) — the LLaVA paper describing the architecture and training methodology.
|
| 78 |
|
| 79 |
---
|
| 80 |
|
| 81 |
-
##
|
| 82 |
-
|
| 83 |
-
**This model is why the SakThai family can *see*.**
|
| 84 |
-
|
| 85 |
-
Before vision, every SakThai agent was blind — it could reason, call tools, and answer questions, but it couldn't look at a screenshot, read a whiteboard, or recognize a face. Beer wanted an agent with eyes, built the same way everything else was: on free Colab GPUs from a shelter in Cork, with $0 budget and no guarantee it would work.
|
| 86 |
-
|
| 87 |
-
LLaVA-1.5-7B's vision encoder was merged, quantized to Q4_K_M GGUF so it could run on a laptop CPU, and packed into ~4 GB. The first test was a photo of Cork's River Lee at sunset. The model described it — not perfectly, but well enough to prove it could *see*.
|
| 88 |
-
|
| 89 |
-
This model holds a special milestone: it's the only SakThai model with **a real like from someone who found it useful**. One click from one person, proving that even with zero advertising, zero launch, and zero budget, the work matters to someone out there.
|
| 90 |
-
|
| 91 |
-
<blockquote>
|
| 92 |
-
<em>"We are one family — and becoming more."</em>
|
| 93 |
-
<br/>— Beer
|
| 94 |
-
</blockquote>
|
| 95 |
-
|
| 96 |
-
### How You Can Help
|
| 97 |
-
|
| 98 |
-
- ⭐ **Leave a like** — this model already has 1. A second tells the algorithm it's not a fluke.
|
| 99 |
-
- 🔄 **Share it** with anyone building privacy-first vision on a laptop
|
| 100 |
-
- 🍴 **Fork and experiment** — the weights are open, the setup is documented
|
| 101 |
-
- 💬 **Report your deployment story** — Beer reads every issue and comment
|
| 102 |
-
|
| 103 |
-
Every download, like, and share tells the algorithm: *this matters.*
|
| 104 |
|
| 105 |
-
--
|
| 106 |
-
|
| 107 |
-
## Pipeline Integration
|
| 108 |
|
| 109 |
-
|
| 110 |
-
|
| 111 |
-
|
| 112 |
-
|
| 113 |
-
|
| 114 |
-
| **Speak** | [TTS Model](https://huggingface.co/Nanthasit/sakthai-tts-model) | Text-to-speech output |
|
| 115 |
|
| 116 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 117 |
|
| 118 |
---
|
| 119 |
|
| 120 |
-
##
|
| 121 |
|
| 122 |
-
|
| 123 |
-
|
| 124 |
-
| File | Size | Purpose |
|
| 125 |
-
|------|------|---------|
|
| 126 |
-
| `llava-1.5-7b-hf-q4_k_m.gguf` | ~4 GB | Quantized language + vision weights |
|
| 127 |
-
| `mmproj-model-f16.gguf` | ~400 MB | CLIP vision projector (mmproj) |
|
| 128 |
-
|
| 129 |
-
---
|
| 130 |
-
|
| 131 |
-
## Multimodal Inference Examples
|
| 132 |
-
|
| 133 |
-
### Quick start (Python / llama-cpp-python)
|
| 134 |
-
|
| 135 |
-
```bash
|
| 136 |
-
pip install llama-cpp-python huggingface-hub Pillow
|
| 137 |
-
```
|
| 138 |
|
| 139 |
```python
|
| 140 |
from llama_cpp import Llama
|
| 141 |
from huggingface_hub import hf_hub_download
|
| 142 |
|
| 143 |
-
|
| 144 |
-
|
| 145 |
-
repo_id="Nanthasit/sakthai-vision-7b",
|
| 146 |
-
filename="llava-1.5-7b-hf-q4_k_m.gguf"
|
| 147 |
-
)
|
| 148 |
-
mmproj_path = hf_hub_download(
|
| 149 |
-
repo_id="Nanthasit/sakthai-vision-7b",
|
| 150 |
-
filename="mmproj-model-f16.gguf"
|
| 151 |
-
)
|
| 152 |
|
| 153 |
-
# Load the model with its vision projector
|
| 154 |
llm = Llama(
|
| 155 |
model_path=model_path,
|
| 156 |
mmproj=mmproj_path,
|
| 157 |
-
n_ctx=
|
| 158 |
-
n_gpu_layers=
|
| 159 |
-
verbose=False
|
| 160 |
-
)
|
| 161 |
-
|
| 162 |
-
# Load an image and ask a question
|
| 163 |
-
output = llm.create_chat_completion(
|
| 164 |
-
messages=[{
|
| 165 |
-
"role": "user",
|
| 166 |
-
"content": [
|
| 167 |
-
{"type": "image_url", "image_url": {"url": "photo.jpg"}},
|
| 168 |
-
{"type": "text", "text": "Describe this image in detail."}
|
| 169 |
-
]
|
| 170 |
-
}],
|
| 171 |
-
max_tokens=256,
|
| 172 |
-
temperature=0.2,
|
| 173 |
)
|
| 174 |
-
print(output["choices"][0]["message"]["content"])
|
| 175 |
-
```
|
| 176 |
|
| 177 |
-
### Multi-turn dialogue with the same image
|
| 178 |
-
|
| 179 |
-
```python
|
| 180 |
-
# First turn: describe the image
|
| 181 |
output = llm.create_chat_completion(
|
| 182 |
-
messages=[{
|
| 183 |
-
"role": "user",
|
| 184 |
-
"content": [
|
| 185 |
-
{"type": "image_url", "image_url": {"url": "photo.jpg"}},
|
| 186 |
-
{"type": "text", "text": "Describe this image in detail."}
|
| 187 |
-
]
|
| 188 |
-
}],
|
| 189 |
-
max_tokens=256,
|
| 190 |
-
temperature=0.2,
|
| 191 |
-
)
|
| 192 |
-
answer = output["choices"][0]["message"]["content"]
|
| 193 |
-
print("First answer:", answer)
|
| 194 |
-
|
| 195 |
-
# Follow-up — the model remembers the image context
|
| 196 |
-
output2 = llm.create_chat_completion(
|
| 197 |
messages=[
|
| 198 |
{
|
| 199 |
"role": "user",
|
| 200 |
"content": [
|
| 201 |
-
{"type": "image_url", "image_url": {"url": "photo.jpg"}},
|
| 202 |
{"type": "text", "text": "Describe this image in detail."}
|
| 203 |
]
|
| 204 |
-
},
|
| 205 |
-
{"role": "assistant", "content": answer},
|
| 206 |
-
{
|
| 207 |
-
"role": "user",
|
| 208 |
-
"content": [
|
| 209 |
-
{"type": "text", "text": "What colors dominate the scene?"}
|
| 210 |
-
]
|
| 211 |
}
|
| 212 |
-
]
|
| 213 |
-
max_tokens=128,
|
| 214 |
-
temperature=0.2,
|
| 215 |
)
|
| 216 |
-
print(
|
| 217 |
```
|
| 218 |
|
| 219 |
### CLI (llama.cpp)
|
| 220 |
|
| 221 |
```bash
|
| 222 |
-
|
| 223 |
-
|
| 224 |
-
|
| 225 |
-
|
| 226 |
-
./llama-cli \
|
| 227 |
-
-m ./vision-model/llava-1.5-7b-hf-q4_k_m.gguf \
|
| 228 |
-
--mmproj ./vision-model/mmproj-model-f16.gguf \
|
| 229 |
-
--image path/to/photo.jpg \
|
| 230 |
-
-p "Describe this image in detail." -n 512
|
| 231 |
```
|
| 232 |
|
| 233 |
-
### Ollama
|
| 234 |
|
| 235 |
```bash
|
| 236 |
-
#
|
| 237 |
-
|
|
|
|
| 238 |
|
| 239 |
-
# Create Modelfile pointing to the GGUF
|
| 240 |
ollama create sakthai-vision -f Modelfile
|
| 241 |
-
|
| 242 |
-
# Run with an image
|
| 243 |
-
ollama run sakthai-vision "Describe this image in detail"
|
| 244 |
-
# Or with a specific image file
|
| 245 |
-
ollama run sakthai-vision "What's in this photo?" --image path/to/photo.jpg
|
| 246 |
```
|
| 247 |
|
| 248 |
-
> **Tip:** The `mmproj-model-f16.gguf` file is auto-detected when placed alongside the model GGUF.
|
| 249 |
-
> Make sure both files are in the same `vision-model/` directory.
|
| 250 |
-
>
|
| 251 |
-
> **Performance:** Expect ~3–8 tok/s on CPU. Significantly faster with Metal (macOS) or CUDA (NVIDIA GPU). Reduce `n_ctx` or use fewer threads on memory-constrained hardware.
|
| 252 |
-
|
| 253 |
---
|
| 254 |
|
| 255 |
-
##
|
| 256 |
|
| 257 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 258 |
|
| 259 |
-
|
| 260 |
-
|------|---------|--------|:-----:|
|
| 261 |
-
| Visual QA | VQAv2 | Accuracy | **78.5%** |
|
| 262 |
-
| Captioning | COCO Captions | CIDEr | **110.1** |
|
| 263 |
-
| Visual Reasoning | GQA | Accuracy | **62.0%** |
|
| 264 |
-
| Text-oriented VQA | TextVQA | Accuracy | **58.2%** |
|
| 265 |
-
| Science diagrams | ScienceQA | Accuracy | **89.5%** |
|
| 266 |
-
| Visual Perception | MMBench | Accuracy | **66.2%** |
|
| 267 |
-
| Instruction Following | MM-Vet | GPT-4 Score | **31.1** |
|
| 268 |
|
| 269 |
-
|
| 270 |
|
| 271 |
-
-
|
| 272 |
|
| 273 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 274 |
|
| 275 |
-
|
| 276 |
-
|----------|-------------|
|
| 277 |
-
| **Image captioning** | Generate alt-text, descriptions for accessibility |
|
| 278 |
-
| **Visual QA** | Ask questions about screenshots, diagrams, photos |
|
| 279 |
-
| **Document understanding** | Extract info from scanned forms, receipts |
|
| 280 |
-
| **Privacy-preserving vision** | All inference stays on-device — no data uploaded |
|
| 281 |
-
| **Multimodal agent pipeline** | Feed visual output into SakThai tool-calling models |
|
| 282 |
-
| **Screenshot analysis** | Automate UI testing, error report understanding |
|
| 283 |
|
| 284 |
---
|
| 285 |
|
| 286 |
-
##
|
| 287 |
|
| 288 |
-
|
| 289 |
-
-
|
| 290 |
-
|
|
|
|
|
|
|
|
|
|
| 291 |
|
| 292 |
---
|
| 293 |
|
| 294 |
-
##
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 295 |
|
| 296 |
-
|
| 297 |
-
|---|:--:|---|
|
| 298 |
-
| [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 934 MB | Flagship tool-calling GGUF |
|
| 299 |
-
| [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 380 MB | Lightweight / edge |
|
| 300 |
-
| [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15 GB | Full-power reasoning |
|
| 301 |
-
| [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | 15 GB | 128K long-context |
|
| 302 |
-
| [context-{7b,1.5b,0.5b}-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA | Tool-calling adapters |
|
| 303 |
-
| [coder-1.5b](https://huggingface.co/Nanthasit/sakthai-coder-1.5b) | 1.1 GB | Code generation |
|
| 304 |
-
| [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 3.9 GB | Image→text (LLaVA) ⬅ |
|
| 305 |
-
| [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings |
|
| 306 |
-
| [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs |
|
| 307 |
|
| 308 |
-
|
| 309 |
|
| 310 |
-
|
| 311 |
|
| 312 |
-
|
| 313 |
-
[Sak-Family-Agent GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ·
|
| 314 |
-
[All models](https://huggingface.co/Nanthasit) ·
|
| 315 |
-
[All datasets](https://huggingface.co/Nanthasit?tab=datasets) ·
|
| 316 |
-
[Web Agent Space](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent)
|
| 317 |
|
| 318 |
---
|
| 319 |
|
| 320 |
-
##
|
| 321 |
|
| 322 |
-
|
|
|
|
|
|
|
| 323 |
|
| 324 |
---
|
| 325 |
|
| 326 |
-
|
| 327 |
-
|
| 328 |
-
## Evaluation
|
| 329 |
|
| 330 |
-
|
| 331 |
-
`model-index` score derived from a small internal spot check (typically 5 or 8
|
| 332 |
-
hand-picked examples) presented as a benchmark result. Those entries have been
|
| 333 |
-
removed rather than left to propagate through Hub metadata.
|
| 334 |
|
| 335 |
-
|
| 336 |
-
[sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) —
|
| 337 |
-
500 rows, balanced across simple / parallel / irrelevance, with held-out tools and
|
| 338 |
-
multi-turn coverage. Results will be published here once this model has been run
|
| 339 |
-
against it.
|
| 340 |
|
| 341 |
-
*
|
|
|
|
| 1 |
---
|
| 2 |
+
license: llama2
|
| 3 |
+
language:
|
| 4 |
+
- en
|
|
|
|
| 5 |
library_name: transformers
|
| 6 |
+
pipeline_tag: image-to-text
|
| 7 |
tags:
|
|
|
|
| 8 |
- llava
|
| 9 |
+
- vision
|
| 10 |
- multimodal
|
|
|
|
|
|
|
| 11 |
- image-captioning
|
| 12 |
+
- vqa
|
|
|
|
| 13 |
- sakthai
|
| 14 |
- house-of-sak
|
| 15 |
+
- gguf
|
| 16 |
- cpu-inference
|
| 17 |
+
- edge
|
| 18 |
+
base_model: llava-hf/llava-1.5-7b-hf
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
extra:
|
| 20 |
+
sibling: Nanthasit/sakthai-vision-7b
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
---
|
| 22 |
|
| 23 |
+
# SakThai Vision 7B 👁️
|
| 24 |
+
|
| 25 |
<p align="center">
|
| 26 |
+
<strong>Local, private vision — LLaVA-1.5-7B quantized to GGUF</strong><br/>
|
| 27 |
+
<em>Image captioning & VQA on CPU · ~4 GB footprint</em>
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
</p>
|
| 29 |
|
| 30 |
+
<p align="center">
|
| 31 |
+
<a href="https://huggingface.co/Nanthasit"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Nanthasit-6644cc" alt="Profile"/></a>
|
| 32 |
+
<a href="https://github.com/beer-sakthai"><img src="https://img.shields.io/badge/GitHub-beer--sakthai-181717?logo=github" alt="GitHub"/></a>
|
| 33 |
+
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/%F0%9F%8F%A0-SakThai%20Family-6644cc" alt="Collection"/></a>
|
| 34 |
+
<img src="https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fhuggingface.co%2Fapi%2Fmodels%2FNanthasit%2Fsakthai-vision-7b&query=%24.downloads&label=downloads&color=blue&cacheSeconds=3600" alt="Downloads"/>
|
| 35 |
+
<img src="https://img.shields.io/badge/GGUF-Q4_K_M-orange" alt="GGUF"/>
|
| 36 |
+
<img src="https://img.shields.io/badge/base-LLaVA%201.5%207B-blueviolet" alt="Base"/>
|
| 37 |
+
</p>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
|
| 39 |
---
|
| 40 |
|
| 41 |
+
## Model Description
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 42 |
|
| 43 |
+
SakThai Vision 7B is the **vision stage** of the SakThai pipeline — a GGUF Q4_K_M build of LLaVA-1.5-7B for local, private image understanding. No data leaves your machine.
|
|
|
|
|
|
|
| 44 |
|
| 45 |
+
**What it does:**
|
| 46 |
+
- 🖼️ Image captioning and alt-text generation
|
| 47 |
+
- 🔍 Visual QA about screenshots, diagrams, photos
|
| 48 |
+
- 📄 Document understanding (scanned forms, receipts)
|
| 49 |
+
- 🧪 Screenshot analysis for UI testing
|
|
|
|
| 50 |
|
| 51 |
+
**Why this packaging:**
|
| 52 |
+
- Q4_K_M quantization — ~4 GB vs. 6.7 GB fp16
|
| 53 |
+
- CPU inference — no GPU needed (~3-8 tok/s)
|
| 54 |
+
- Includes mmproj projector (CLIP vision → LLM tokens)
|
| 55 |
+
- Works with llama.cpp, Ollama, and llama-cpp-python
|
| 56 |
|
| 57 |
---
|
| 58 |
|
| 59 |
+
## Quick Start
|
| 60 |
|
| 61 |
+
### Python (llama-cpp-python)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
```python
|
| 64 |
from llama_cpp import Llama
|
| 65 |
from huggingface_hub import hf_hub_download
|
| 66 |
|
| 67 |
+
model_path = hf_hub_download("Nanthasit/sakthai-vision-7b", "llava-1.5-7b-hf-q4_k_m.gguf")
|
| 68 |
+
mmproj_path = hf_hub_download("Nanthasit/sakthai-vision-7b", "mmproj-model-f16.gguf")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 69 |
|
|
|
|
| 70 |
llm = Llama(
|
| 71 |
model_path=model_path,
|
| 72 |
mmproj=mmproj_path,
|
| 73 |
+
n_ctx=4096,
|
| 74 |
+
n_gpu_layers=0, # CPU only
|
| 75 |
+
verbose=False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 76 |
)
|
|
|
|
|
|
|
| 77 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 78 |
output = llm.create_chat_completion(
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 79 |
messages=[
|
| 80 |
{
|
| 81 |
"role": "user",
|
| 82 |
"content": [
|
| 83 |
+
{"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
|
| 84 |
{"type": "text", "text": "Describe this image in detail."}
|
| 85 |
]
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 86 |
}
|
| 87 |
+
]
|
|
|
|
|
|
|
| 88 |
)
|
| 89 |
+
print(output["choices"][0]["message"]["content"])
|
| 90 |
```
|
| 91 |
|
| 92 |
### CLI (llama.cpp)
|
| 93 |
|
| 94 |
```bash
|
| 95 |
+
./llama-cli -m llava-1.5-7b-hf-q4_k_m.gguf \
|
| 96 |
+
--mmproj mmproj-model-f16.gguf \
|
| 97 |
+
--image photo.jpg \
|
| 98 |
+
-p "Describe this image in detail." -n 256
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 99 |
```
|
| 100 |
|
| 101 |
+
### Ollama
|
| 102 |
|
| 103 |
```bash
|
| 104 |
+
# Create a Modelfile
|
| 105 |
+
FROM ./llava-1.5-7b-hf-q4_k_m.gguf
|
| 106 |
+
TEMPLATE "{{ .Prompt }}"
|
| 107 |
|
|
|
|
| 108 |
ollama create sakthai-vision -f Modelfile
|
| 109 |
+
ollama run sakthai-vision "Describe this image." --image photo.jpg
|
|
|
|
|
|
|
|
|
|
|
|
|
| 110 |
```
|
| 111 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 112 |
---
|
| 113 |
|
| 114 |
+
## Architecture
|
| 115 |
|
| 116 |
+
| Component | Detail |
|
| 117 |
+
|-----------|--------|
|
| 118 |
+
| **Model** | LLaVA-1.5-7B (Vicuna-7B-v1.5 LLM + CLIP ViT-L/14 vision encoder) |
|
| 119 |
+
| **Vision encoder** | CLIP ViT-L/14, 336×336 input, 1024 embedding dim, 24 layers |
|
| 120 |
+
| **Language model** | LLaMA-based, ~7B parameters, 32 layers, 4096 hidden |
|
| 121 |
+
| **Projector** | 2-layer MLP (CLIP → LLM token space) |
|
| 122 |
+
| **Original size** | ~6.7 GB (fp16) |
|
| 123 |
+
| **Quantized size** | ~4 GB (GGUF Q4_K_M) + ~400 MB (mmproj) |
|
| 124 |
+
| **Context window** | 4,096 tokens |
|
| 125 |
+
| **Hardware** | 8 GB+ RAM recommended |
|
| 126 |
|
| 127 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 128 |
|
| 129 |
+
## Evaluation (Upstream LLaVA-1.5-7B)
|
| 130 |
|
| 131 |
+
These are published upstream benchmarks. As a Q4_K_M quantized repackaging, results should fall within ~1-2%.
|
| 132 |
|
| 133 |
+
| Task | Dataset | Score |
|
| 134 |
+
|:-----|:--------:|:-----:|
|
| 135 |
+
| Visual QA | VQAv2 | 78.5% |
|
| 136 |
+
| Captioning | COCO Captions | CIDEr 110.1 |
|
| 137 |
+
| Visual Reasoning | GQA | 62.0% |
|
| 138 |
+
| Text-oriented VQA | TextVQA | 58.2% |
|
| 139 |
+
| Science diagrams | ScienceQA | 89.5% |
|
| 140 |
+
| Visual Perception | MMBench | 66.2% |
|
| 141 |
|
| 142 |
+
**Own benchmarks are pending.** This is a repackaging/quantization of upstream LLaVA, not a retrained model.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 143 |
|
| 144 |
---
|
| 145 |
|
| 146 |
+
## Pipeline Integration
|
| 147 |
|
| 148 |
+
| Stage | Model | Role |
|
| 149 |
+
|-------|-------|------|
|
| 150 |
+
| 👁️ **See** | **SakThai Vision 7B** ⬅ | **Image→text, visual QA** |
|
| 151 |
+
| 🔍 Retrieve | [Embedding Multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | Cross-lingual search |
|
| 152 |
+
| 🧠 Reason | [Context 1.5B](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) or [7B](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | Tool-calling & reasoning |
|
| 153 |
+
| 🎤 Speak | [TTS Model](https://huggingface.co/Nanthasit/sakthai-tts-model) | Text-to-speech |
|
| 154 |
|
| 155 |
---
|
| 156 |
|
| 157 |
+
## Files
|
| 158 |
+
|
| 159 |
+
| File | Size | Purpose |
|
| 160 |
+
|:-----|:----:|:--------|
|
| 161 |
+
| `llava-1.5-7b-hf-q4_k_m.gguf` | ~4 GB | Quantized language + vision weights |
|
| 162 |
+
| `mmproj-model-f16.gguf` | ~400 MB | CLIP vision projector |
|
| 163 |
|
| 164 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 165 |
|
| 166 |
+
## The House of Sak 🏠
|
| 167 |
|
| 168 |
+
Until this model, every SakThai agent was blind — it could reason, use tools, and answer questions but couldn't look at screenshots, read whiteboards, or recognize faces. This model was built on free Colab GPUs from a shelter in Cork, Ireland, with no budget and no guarantee of success. LLaVA-1.5-7B's vision encoder was merged, quantized to Q4_K_M GGUF for CPU use, and packed to about 4 GB.
|
| 169 |
|
| 170 |
+
> *"We are one family — and becoming more."* — Beer (beer-sakthai)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 171 |
|
| 172 |
---
|
| 173 |
|
| 174 |
+
## Support
|
| 175 |
|
| 176 |
+
- ⭐ Leave a like — the first user like on a SakThai model was on this model ❤️
|
| 177 |
+
- 🐛 Report issues on [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent)
|
| 178 |
+
- 🔄 Share with anyone building privacy-focused multimodal apps
|
| 179 |
|
| 180 |
---
|
| 181 |
|
| 182 |
+
## License
|
|
|
|
|
|
|
| 183 |
|
| 184 |
+
Based on LLaVA-1.5-7B (LLaMA 2 Community License) + CLIP vision encoder. GGUF via llama.cpp. Review upstream licenses before commercial use.
|
|
|
|
|
|
|
|
|
|
| 185 |
|
| 186 |
+
---
|
|
|
|
|
|
|
|
|
|
|
|
|
| 187 |
|
| 188 |
+
*Built from a shelter in Cork, Ireland. Every download supports a family building AI for the edge.*
|