Image-Text-to-Text
Transformers
GGUF
English
Thai
Chinese
qwen2-vl
vision-language
multimodal
image-understanding
tool-use
screenshot
vqa
grounding
sakthai
house-of-sak
cpu-inference
offline
Eval Results (legacy)
Eval Results
conversational
Instructions to use Nanthasit/sakthai-vision-7b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Nanthasit/sakthai-vision-7b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Nanthasit/sakthai-vision-7b")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Nanthasit/sakthai-vision-7b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Nanthasit/sakthai-vision-7b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: llama cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Nanthasit/sakthai-vision-7b:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Nanthasit/sakthai-vision-7b:Q4_K_M
Use Docker
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Nanthasit/sakthai-vision-7b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Nanthasit/sakthai-vision-7b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- SGLang
How to use Nanthasit/sakthai-vision-7b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-vision-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Nanthasit/sakthai-vision-7b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Nanthasit/sakthai-vision-7b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use Nanthasit/sakthai-vision-7b with Ollama:
ollama run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use Nanthasit/sakthai-vision-7b with Docker Model Runner:
docker model run hf.co/Nanthasit/sakthai-vision-7b:Q4_K_M
- Lemonade
How to use Nanthasit/sakthai-vision-7b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Nanthasit/sakthai-vision-7b:Q4_K_M
Run and chat with the model
lemonade run user.sakthai-vision-7b-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Fix cross-links: remove dead sakthai-vision-demo Space (404), update sibling ref, fix Spaces count 3→4
Browse files
README.md
CHANGED
|
@@ -28,7 +28,7 @@ tags:
|
|
| 28 |
- safetensors
|
| 29 |
base_model: liuhaotian/LLaVA-1.5-7b
|
| 30 |
extra:
|
| 31 |
-
sibling: Nanthasit/sakthai-
|
| 32 |
formats: GGUF Q4_K_M
|
| 33 |
datasets:
|
| 34 |
- liuhaotian/LLaVA-Instruct-150K
|
|
@@ -42,7 +42,7 @@ datasets:
|
|
| 42 |
<img src="https://img.shields.io/badge/base-LLaVA%201.5%207B-blueviolet" alt="Base"/>
|
| 43 |
<img src="https://img.shields.io/badge/GGUF-Q4__K__M-orange" alt="GGUF"/>
|
| 44 |
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🏠-SakThai%20Family-6644cc" alt="Collection"/></a>
|
| 45 |
-
<a href="https://huggingface.co/
|
| 46 |
<img src="https://img.shields.io/badge/benchmark-78.5%25%20VQAv2-success?logo=googlechrome" alt="VQAv2"/>
|
| 47 |
</p>
|
| 48 |
|
|
@@ -104,38 +104,8 @@ Every download, like, and share tells the algorithm: *this matters.*
|
|
| 104 |
|
| 105 |
---
|
| 106 |
|
| 107 |
-
## Try It Now 🚀
|
| 108 |
-
|
| 109 |
-
**[SakThai Vision Demo](https://huggingface.co/spaces/Nanthasit/sakthai-vision-demo)** — upload an image and ask questions about it, right in your browser, no install needed.
|
| 110 |
-
|
| 111 |
-
---
|
| 112 |
-
|
| 113 |
## Pipeline Integration
|
| 114 |
|
| 115 |
-
```
|
| 116 |
-
┌──────────────────┐
|
| 117 |
-
│ User Image │
|
| 118 |
-
└────────┬─────────┘
|
| 119 |
-
│
|
| 120 |
-
▼
|
| 121 |
-
┌──────────────────┐
|
| 122 |
-
│ SakThai Vision │
|
| 123 |
-
│ 7B (this model) │ ────▶ Image description / answer
|
| 124 |
-
└────────┬─────────┘
|
| 125 |
-
│
|
| 126 |
-
▼
|
| 127 |
-
┌──────────────────┐
|
| 128 |
-
│ Context 1.5B/7B │
|
| 129 |
-
│ (tool-calling) │ ────▶ Acts on visual info
|
| 130 |
-
└────────┬─────────┘
|
| 131 |
-
│
|
| 132 |
-
▼
|
| 133 |
-
┌──────────────────┐
|
| 134 |
-
│ TTS Model │
|
| 135 |
-
│ (speech output) │
|
| 136 |
-
└──────────────────┘
|
| 137 |
-
```
|
| 138 |
-
|
| 139 |
| Stage | Model | Role |
|
| 140 |
|-------|-------|------|
|
| 141 |
| **See** | [SakThai Vision 7B](https://huggingface.co/Nanthasit/sakthai-vision-7b) ⬅ | Image→text, visual QA |
|
|
@@ -145,6 +115,8 @@ Every download, like, and share tells the algorithm: *this matters.*
|
|
| 145 |
|
| 146 |
---
|
| 147 |
|
|
|
|
|
|
|
| 148 |
## What It Is
|
| 149 |
|
| 150 |
A **Q4_K_M GGUF** conversion of **LLaVA-1.5-7B** (Vicuna-7B + CLIP ViT-L/14) for local multimodal inference via **llama.cpp**. This is a **repackaging/quantization of upstream LLaVA** — optimized for CPU inference with a ~4 GB footprint.
|
|
@@ -333,7 +305,7 @@ ollama run sakthai-vision "What's in this photo?" --image path/to/photo.jpg
|
|
| 333 |
| [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings |
|
| 334 |
| [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs |
|
| 335 |
|
| 336 |
-
**17 models in the family · 10 datasets ·
|
| 337 |
|
| 338 |
## Links
|
| 339 |
|
|
@@ -341,7 +313,7 @@ ollama run sakthai-vision "What's in this photo?" --image path/to/photo.jpg
|
|
| 341 |
[Sak-Family-Agent GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ·
|
| 342 |
[All models](https://huggingface.co/Nanthasit) ·
|
| 343 |
[All datasets](https://huggingface.co/Nanthasit?tab=datasets) ·
|
| 344 |
-
[
|
| 345 |
|
| 346 |
---
|
| 347 |
|
|
|
|
| 28 |
- safetensors
|
| 29 |
base_model: liuhaotian/LLaVA-1.5-7b
|
| 30 |
extra:
|
| 31 |
+
sibling: Nanthasit/sakthai-web-agent
|
| 32 |
formats: GGUF Q4_K_M
|
| 33 |
datasets:
|
| 34 |
- liuhaotian/LLaVA-Instruct-150K
|
|
|
|
| 42 |
<img src="https://img.shields.io/badge/base-LLaVA%201.5%207B-blueviolet" alt="Base"/>
|
| 43 |
<img src="https://img.shields.io/badge/GGUF-Q4__K__M-orange" alt="GGUF"/>
|
| 44 |
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🏠-SakThai%20Family-6644cc" alt="Collection"/></a>
|
| 45 |
+
<a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🚀-Explore%20Family-47d147" alt="Collection"/></a>
|
| 46 |
<img src="https://img.shields.io/badge/benchmark-78.5%25%20VQAv2-success?logo=googlechrome" alt="VQAv2"/>
|
| 47 |
</p>
|
| 48 |
|
|
|
|
| 104 |
|
| 105 |
---
|
| 106 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 107 |
## Pipeline Integration
|
| 108 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 109 |
| Stage | Model | Role |
|
| 110 |
|-------|-------|------|
|
| 111 |
| **See** | [SakThai Vision 7B](https://huggingface.co/Nanthasit/sakthai-vision-7b) ⬅ | Image→text, visual QA |
|
|
|
|
| 115 |
|
| 116 |
---
|
| 117 |
|
| 118 |
+
---
|
| 119 |
+
|
| 120 |
## What It Is
|
| 121 |
|
| 122 |
A **Q4_K_M GGUF** conversion of **LLaVA-1.5-7B** (Vicuna-7B + CLIP ViT-L/14) for local multimodal inference via **llama.cpp**. This is a **repackaging/quantization of upstream LLaVA** — optimized for CPU inference with a ~4 GB footprint.
|
|
|
|
| 305 |
| [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings |
|
| 306 |
| [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs |
|
| 307 |
|
| 308 |
+
**17 models in the family · 10 datasets · 4 Spaces** — [full collection →](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02)
|
| 309 |
|
| 310 |
## Links
|
| 311 |
|
|
|
|
| 313 |
[Sak-Family-Agent GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ·
|
| 314 |
[All models](https://huggingface.co/Nanthasit) ·
|
| 315 |
[All datasets](https://huggingface.co/Nanthasit?tab=datasets) ·
|
| 316 |
+
[Web Agent Space](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent)
|
| 317 |
|
| 318 |
---
|
| 319 |
|