Nanthasit commited on
Commit
810fd3d
·
verified ·
1 Parent(s): 8fd6e28

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +101 -254
README.md CHANGED
@@ -1,341 +1,188 @@
1
  ---
2
- license: other
3
- license_name: llama2-community
4
- license_link: https://ai.meta.com/llama/license/
5
- pipeline_tag: image-to-text
6
  library_name: transformers
 
7
  tags:
8
- - vision
9
  - llava
 
10
  - multimodal
11
- - image-to-text
12
- - visual-question-answering
13
  - image-captioning
14
- - gguf
15
- - llama-cpp
16
  - sakthai
17
  - house-of-sak
18
- - edge
19
  - cpu-inference
20
- - local-ai
21
- - offline
22
- - privacy
23
- - quantized
24
- - transformers
25
- - clip
26
- - vicuna
27
- - finetune
28
- - safetensors
29
- base_model: liuhaotian/LLaVA-1.5-7b
30
  extra:
31
- sibling: Nanthasit/sakthai-web-agent
32
- formats: GGUF Q4_K_M
33
- datasets:
34
- - liuhaotian/LLaVA-Instruct-150K
35
- - liuhaotian/LLaVA-Pretrain
36
  ---
37
 
38
- <h1 align="center">SakThai Vision 7B 👁️</h1>
39
- <p align="center"><em>LLaVA-1.5-7B, quantized to GGUF · image captioning & VQA on CPU · local & private</em></p>
40
  <p align="center">
41
- <img src="https://img.shields.io/badge/dynamic/json?url=https%3A//huggingface.co/api/models/Nanthasit/sakthai-vision-7b&query=%24.downloads&label=downloads&color=blue&cacheSeconds=3600" alt="Downloads"/>
42
- <img src="https://img.shields.io/badge/base-LLaVA%201.5%207B-blueviolet" alt="Base"/>
43
- <img src="https://img.shields.io/badge/GGUF-Q4__K__M-orange" alt="GGUF"/>
44
- <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🏠-SakThai%20Family-6644cc" alt="Collection"/></a>
45
- <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/🚀-Explore%20Family-47d147" alt="Collection"/></a>
46
- <img src="https://img.shields.io/badge/benchmark-78.5%25%20VQAv2-success?logo=googlechrome" alt="VQAv2"/>
47
  </p>
48
 
49
- > The **vision stage** of the SakThai multimodal pipeline — a GGUF build of **LLaVA-1.5-7B** for local,
50
- > private image understanding. Part of the [House of Sak](https://huggingface.co/Nanthasit).
51
- > No data leaves your machine.
52
- >
53
- > **🏆 First SakThai model with a real like** — someone found this useful. [Leave yours →](#how-you-can-help)
54
-
55
- ---
56
-
57
- ## Architecture
58
-
59
- | Component | Detail |
60
- |-----------|--------|
61
- | **Model** | LLaVA-1.5-7B ([paper](https://arxiv.org/abs/2310.03744)) — Vicuna-7B-v1.5 LLM + CLIP ViT-L/14 vision encoder |
62
- | **Vision encoder** | CLIP ViT-L/14, 336×336 input, 1024 embedding dim, 24 layers, 16 heads, ~427M params |
63
- | **Language model** | LLaMA-based, ~7B parameters, 32 layers, 4096 hidden dim, 32 heads |
64
- | **Projector** | 2-layer MLP (CLIP embedding → LLM token space), 4096 hidden dim |
65
- | **Context window** | 4096 tokens |
66
- | **Attention** | Causal self-attention (LLM) + bidirectional self-attention (vision encoder) |
67
- | **Quantization** | GGUF Q4_K_M (4-bit group quantization, 256 group size) |
68
- | **File size** | ~4 GB (LLM) + ~400 MB (mmproj) |
69
- | **Original size** | ~6.7 GB (unquantized fp16) |
70
- | **Chat template** | Vicuna-style with `<image>` placeholder |
71
- | **License** | LLaMA 2 Community License |
72
-
73
- LLaVA (Large Language and Vision Assistant) connects a pre-trained CLIP vision encoder to a Vicuna LLM via a lightweight projection MLP. The vision encoder processes images into embeddings, the projector maps them into the LLM's token space, and the LLM generates text conditioned on both the visual and textual input. This architecture enables multi-turn visual dialogue: once an image is encoded, you can ask follow-up questions without reprocessing the visual input.
74
-
75
- This quantized GGUF build reduces the footprint from ~6.7 GB (fp16) to ~4 GB while retaining ~98–99% of the original accuracy.
76
-
77
- > **Reference:** Liu et al., *"Visual Instruction Tuning"* (NeurIPS 2023). [arXiv:2310.03744](https://arxiv.org/abs/2310.03744) — the LLaVA paper describing the architecture and training methodology.
78
 
79
  ---
80
 
81
- ## The Story Behind It
82
-
83
- **This model is why the SakThai family can *see*.**
84
-
85
- Before vision, every SakThai agent was blind — it could reason, call tools, and answer questions, but it couldn't look at a screenshot, read a whiteboard, or recognize a face. Beer wanted an agent with eyes, built the same way everything else was: on free Colab GPUs from a shelter in Cork, with $0 budget and no guarantee it would work.
86
-
87
- LLaVA-1.5-7B's vision encoder was merged, quantized to Q4_K_M GGUF so it could run on a laptop CPU, and packed into ~4 GB. The first test was a photo of Cork's River Lee at sunset. The model described it — not perfectly, but well enough to prove it could *see*.
88
-
89
- This model holds a special milestone: it's the only SakThai model with **a real like from someone who found it useful**. One click from one person, proving that even with zero advertising, zero launch, and zero budget, the work matters to someone out there.
90
-
91
- <blockquote>
92
- <em>"We are one family — and becoming more."</em>
93
- <br/>— Beer
94
- </blockquote>
95
-
96
- ### How You Can Help
97
-
98
- - ⭐ **Leave a like** — this model already has 1. A second tells the algorithm it's not a fluke.
99
- - 🔄 **Share it** with anyone building privacy-first vision on a laptop
100
- - 🍴 **Fork and experiment** — the weights are open, the setup is documented
101
- - 💬 **Report your deployment story** — Beer reads every issue and comment
102
-
103
- Every download, like, and share tells the algorithm: *this matters.*
104
 
105
- ---
106
-
107
- ## Pipeline Integration
108
 
109
- | Stage | Model | Role |
110
- |-------|-------|------|
111
- | **See** | [SakThai Vision 7B](https://huggingface.co/Nanthasit/sakthai-vision-7b) ⬅ | Image→text, visual QA |
112
- | **Embed** | [SakThai Embedding](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | Convert queries/docs to 384d vectors |
113
- | **Reason** | [1.5B-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) or [7B-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | Tool calling + response generation |
114
- | **Speak** | [TTS Model](https://huggingface.co/Nanthasit/sakthai-tts-model) | Text-to-speech output |
115
 
116
- ---
 
 
 
 
117
 
118
  ---
119
 
120
- ## What It Is
121
 
122
- A **Q4_K_M GGUF** conversion of **LLaVA-1.5-7B** (Vicuna-7B + CLIP ViT-L/14) for local multimodal inference via **llama.cpp**. This is a **repackaging/quantization of upstream LLaVA** — optimized for CPU inference with a ~4 GB footprint.
123
-
124
- | File | Size | Purpose |
125
- |------|------|---------|
126
- | `llava-1.5-7b-hf-q4_k_m.gguf` | ~4 GB | Quantized language + vision weights |
127
- | `mmproj-model-f16.gguf` | ~400 MB | CLIP vision projector (mmproj) |
128
-
129
- ---
130
-
131
- ## Multimodal Inference Examples
132
-
133
- ### Quick start (Python / llama-cpp-python)
134
-
135
- ```bash
136
- pip install llama-cpp-python huggingface-hub Pillow
137
- ```
138
 
139
  ```python
140
  from llama_cpp import Llama
141
  from huggingface_hub import hf_hub_download
142
 
143
- # Download model files from Hugging Face
144
- model_path = hf_hub_download(
145
- repo_id="Nanthasit/sakthai-vision-7b",
146
- filename="llava-1.5-7b-hf-q4_k_m.gguf"
147
- )
148
- mmproj_path = hf_hub_download(
149
- repo_id="Nanthasit/sakthai-vision-7b",
150
- filename="mmproj-model-f16.gguf"
151
- )
152
 
153
- # Load the model with its vision projector
154
  llm = Llama(
155
  model_path=model_path,
156
  mmproj=mmproj_path,
157
- n_ctx=2048, # Context window
158
- n_gpu_layers=-1, # Offload to GPU if available (-1 = all)
159
- verbose=False,
160
- )
161
-
162
- # Load an image and ask a question
163
- output = llm.create_chat_completion(
164
- messages=[{
165
- "role": "user",
166
- "content": [
167
- {"type": "image_url", "image_url": {"url": "photo.jpg"}},
168
- {"type": "text", "text": "Describe this image in detail."}
169
- ]
170
- }],
171
- max_tokens=256,
172
- temperature=0.2,
173
  )
174
- print(output["choices"][0]["message"]["content"])
175
- ```
176
 
177
- ### Multi-turn dialogue with the same image
178
-
179
- ```python
180
- # First turn: describe the image
181
  output = llm.create_chat_completion(
182
- messages=[{
183
- "role": "user",
184
- "content": [
185
- {"type": "image_url", "image_url": {"url": "photo.jpg"}},
186
- {"type": "text", "text": "Describe this image in detail."}
187
- ]
188
- }],
189
- max_tokens=256,
190
- temperature=0.2,
191
- )
192
- answer = output["choices"][0]["message"]["content"]
193
- print("First answer:", answer)
194
-
195
- # Follow-up — the model remembers the image context
196
- output2 = llm.create_chat_completion(
197
  messages=[
198
  {
199
  "role": "user",
200
  "content": [
201
- {"type": "image_url", "image_url": {"url": "photo.jpg"}},
202
  {"type": "text", "text": "Describe this image in detail."}
203
  ]
204
- },
205
- {"role": "assistant", "content": answer},
206
- {
207
- "role": "user",
208
- "content": [
209
- {"type": "text", "text": "What colors dominate the scene?"}
210
- ]
211
  }
212
- ],
213
- max_tokens=128,
214
- temperature=0.2,
215
  )
216
- print("Follow-up answer:", output2["choices"][0]["message"]["content"])
217
  ```
218
 
219
  ### CLI (llama.cpp)
220
 
221
  ```bash
222
- # Download the model files
223
- huggingface-cli download Nanthasit/sakthai-vision-7b --local-dir ./vision-model
224
-
225
- # Run inference — load image, ask question
226
- ./llama-cli \
227
- -m ./vision-model/llava-1.5-7b-hf-q4_k_m.gguf \
228
- --mmproj ./vision-model/mmproj-model-f16.gguf \
229
- --image path/to/photo.jpg \
230
- -p "Describe this image in detail." -n 512
231
  ```
232
 
233
- ### Ollama (local)
234
 
235
  ```bash
236
- # Download model files
237
- huggingface-cli download Nanthasit/sakthai-vision-7b --local-dir ./vision-model
 
238
 
239
- # Create Modelfile pointing to the GGUF
240
  ollama create sakthai-vision -f Modelfile
241
-
242
- # Run with an image
243
- ollama run sakthai-vision "Describe this image in detail"
244
- # Or with a specific image file
245
- ollama run sakthai-vision "What's in this photo?" --image path/to/photo.jpg
246
  ```
247
 
248
- > **Tip:** The `mmproj-model-f16.gguf` file is auto-detected when placed alongside the model GGUF.
249
- > Make sure both files are in the same `vision-model/` directory.
250
- >
251
- > **Performance:** Expect ~3–8 tok/s on CPU. Significantly faster with Metal (macOS) or CUDA (NVIDIA GPU). Reduce `n_ctx` or use fewer threads on memory-constrained hardware.
252
-
253
  ---
254
 
255
- ## Benchmarks (Upstream LLaVA-1.5-7B)
256
 
257
- > These are the **published upstream benchmarks** for LLaVA-1.5-7B. As a quantized repackaging, expect results within ~1–2% of these values on most tasks.
 
 
 
 
 
 
 
 
 
258
 
259
- | Task | Dataset | Metric | Score |
260
- |------|---------|--------|:-----:|
261
- | Visual QA | VQAv2 | Accuracy | **78.5%** |
262
- | Captioning | COCO Captions | CIDEr | **110.1** |
263
- | Visual Reasoning | GQA | Accuracy | **62.0%** |
264
- | Text-oriented VQA | TextVQA | Accuracy | **58.2%** |
265
- | Science diagrams | ScienceQA | Accuracy | **89.5%** |
266
- | Visual Perception | MMBench | Accuracy | **66.2%** |
267
- | Instruction Following | MM-Vet | GPT-4 Score | **31.1** |
268
 
269
- > **Note:** GGUF Q4_K_M quantization may cause minor degradation vs. fp16. Own run-specific benchmarks pending publication.
270
 
271
- ---
272
 
273
- ## Use Cases
 
 
 
 
 
 
 
274
 
275
- | Use Case | Description |
276
- |----------|-------------|
277
- | **Image captioning** | Generate alt-text, descriptions for accessibility |
278
- | **Visual QA** | Ask questions about screenshots, diagrams, photos |
279
- | **Document understanding** | Extract info from scanned forms, receipts |
280
- | **Privacy-preserving vision** | All inference stays on-device — no data uploaded |
281
- | **Multimodal agent pipeline** | Feed visual output into SakThai tool-calling models |
282
- | **Screenshot analysis** | Automate UI testing, error report understanding |
283
 
284
  ---
285
 
286
- ## Intended Use & Limits
287
 
288
- - **✅ Great for:** Image captioning, visual QA, screenshot/diagram analysis, local privacy-sensitive inference, prototyping multimodal agents
289
- - **⚠️ Limits:** Requires llama.cpp with multimodal support — not directly usable with plain Transformers. CLIP native resolution is 336×336 (images are resized). Can hallucinate on ambiguous images. Primarily English.
290
- - **💻 Hardware:** 8 GB+ RAM recommended. ~3–8 tok/s on CPU, faster with Metal/CUDA.
 
 
 
291
 
292
  ---
293
 
294
- ## SakThai model family
 
 
 
 
 
295
 
296
- | Model | Size | Role |
297
- |---|:--:|---|
298
- | [context-1.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) | 934 MB | Flagship tool-calling GGUF |
299
- | [context-0.5b-merged](https://huggingface.co/Nanthasit/sakthai-context-0.5b-merged) | 380 MB | Lightweight / edge |
300
- | [context-7b-merged](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | 15 GB | Full-power reasoning |
301
- | [context-7b-128k](https://huggingface.co/Nanthasit/sakthai-context-7b-128k) | 15 GB | 128K long-context |
302
- | [context-{7b,1.5b,0.5b}-tools](https://huggingface.co/Nanthasit/sakthai-context-1.5b-tools) | LoRA | Tool-calling adapters |
303
- | [coder-1.5b](https://huggingface.co/Nanthasit/sakthai-coder-1.5b) | 1.1 GB | Code generation |
304
- | [vision-7b](https://huggingface.co/Nanthasit/sakthai-vision-7b) | 3.9 GB | Image→text (LLaVA) ⬅ |
305
- | [embedding-multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | 80 MB | Cross-lingual embeddings |
306
- | [tts-model](https://huggingface.co/Nanthasit/sakthai-tts-model) | 141 MB | Text-to-speech, 15 langs |
307
 
308
- **20 models in the family · 10 datasets · 4 Spaces** — [full collection →](https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02)
309
 
310
- ## Links
311
 
312
- [House of Sak](https://house-of-sak.vercel.app) ·
313
- [Sak-Family-Agent GitHub](https://github.com/beer-sakthai/Sak-Family-Agent) ·
314
- [All models](https://huggingface.co/Nanthasit) ·
315
- [All datasets](https://huggingface.co/Nanthasit?tab=datasets) ·
316
- [Web Agent Space](https://huggingface.co/spaces/Nanthasit/sakthai-web-agent)
317
 
318
  ---
319
 
320
- ## License
321
 
322
- Based on LLaVA-1.5-7B (LLaMA 2 Community License) with a CLIP vision encoder; GGUF via llama.cpp. Review the upstream licenses before commercial use.
 
 
323
 
324
  ---
325
 
326
- *Built from a shelter in Cork, Ireland. Every download supports a family building AI for the edge.*
327
-
328
- ## Evaluation
329
 
330
- **Not independently benchmarked.** Earlier versions of this card carried a
331
- `model-index` score derived from a small internal spot check (typically 5 or 8
332
- hand-picked examples) presented as a benchmark result. Those entries have been
333
- removed rather than left to propagate through Hub metadata.
334
 
335
- For tool-calling models in this family, the benchmark to use is
336
- [sakthai-bench-v2](https://huggingface.co/datasets/Nanthasit/sakthai-bench-v2) —
337
- 500 rows, balanced across simple / parallel / irrelevance, with held-out tools and
338
- multi-turn coverage. Results will be published here once this model has been run
339
- against it.
340
 
341
- *"We are one family — and becoming more."* 🏠
 
1
  ---
2
+ license: llama2
3
+ language:
4
+ - en
 
5
  library_name: transformers
6
+ pipeline_tag: image-to-text
7
  tags:
 
8
  - llava
9
+ - vision
10
  - multimodal
 
 
11
  - image-captioning
12
+ - vqa
 
13
  - sakthai
14
  - house-of-sak
15
+ - gguf
16
  - cpu-inference
17
+ - edge
18
+ base_model: llava-hf/llava-1.5-7b-hf
 
 
 
 
 
 
 
 
19
  extra:
20
+ sibling: Nanthasit/sakthai-vision-7b
 
 
 
 
21
  ---
22
 
23
+ # SakThai Vision 7B 👁️
24
+
25
  <p align="center">
26
+ <strong>Local, private vision — LLaVA-1.5-7B quantized to GGUF</strong><br/>
27
+ <em>Image captioning & VQA on CPU · ~4 GB footprint</em>
 
 
 
 
28
  </p>
29
 
30
+ <p align="center">
31
+ <a href="https://huggingface.co/Nanthasit"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Nanthasit-6644cc" alt="Profile"/></a>
32
+ <a href="https://github.com/beer-sakthai"><img src="https://img.shields.io/badge/GitHub-beer--sakthai-181717?logo=github" alt="GitHub"/></a>
33
+ <a href="https://huggingface.co/collections/Nanthasit/sakthai-model-family-6a64745450b12d421c1f9f02"><img src="https://img.shields.io/badge/%F0%9F%8F%A0-SakThai%20Family-6644cc" alt="Collection"/></a>
34
+ <img src="https://img.shields.io/badge/dynamic/json?url=https%3A%2F%2Fhuggingface.co%2Fapi%2Fmodels%2FNanthasit%2Fsakthai-vision-7b&query=%24.downloads&label=downloads&color=blue&cacheSeconds=3600" alt="Downloads"/>
35
+ <img src="https://img.shields.io/badge/GGUF-Q4_K_M-orange" alt="GGUF"/>
36
+ <img src="https://img.shields.io/badge/base-LLaVA%201.5%207B-blueviolet" alt="Base"/>
37
+ </p>
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38
 
39
  ---
40
 
41
+ ## Model Description
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42
 
43
+ SakThai Vision 7B is the **vision stage** of the SakThai pipeline — a GGUF Q4_K_M build of LLaVA-1.5-7B for local, private image understanding. No data leaves your machine.
 
 
44
 
45
+ **What it does:**
46
+ - 🖼️ Image captioning and alt-text generation
47
+ - 🔍 Visual QA about screenshots, diagrams, photos
48
+ - 📄 Document understanding (scanned forms, receipts)
49
+ - 🧪 Screenshot analysis for UI testing
 
50
 
51
+ **Why this packaging:**
52
+ - Q4_K_M quantization — ~4 GB vs. 6.7 GB fp16
53
+ - CPU inference — no GPU needed (~3-8 tok/s)
54
+ - Includes mmproj projector (CLIP vision → LLM tokens)
55
+ - Works with llama.cpp, Ollama, and llama-cpp-python
56
 
57
  ---
58
 
59
+ ## Quick Start
60
 
61
+ ### Python (llama-cpp-python)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
 
63
  ```python
64
  from llama_cpp import Llama
65
  from huggingface_hub import hf_hub_download
66
 
67
+ model_path = hf_hub_download("Nanthasit/sakthai-vision-7b", "llava-1.5-7b-hf-q4_k_m.gguf")
68
+ mmproj_path = hf_hub_download("Nanthasit/sakthai-vision-7b", "mmproj-model-f16.gguf")
 
 
 
 
 
 
 
69
 
 
70
  llm = Llama(
71
  model_path=model_path,
72
  mmproj=mmproj_path,
73
+ n_ctx=4096,
74
+ n_gpu_layers=0, # CPU only
75
+ verbose=False
 
 
 
 
 
 
 
 
 
 
 
 
 
76
  )
 
 
77
 
 
 
 
 
78
  output = llm.create_chat_completion(
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
79
  messages=[
80
  {
81
  "role": "user",
82
  "content": [
83
+ {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}},
84
  {"type": "text", "text": "Describe this image in detail."}
85
  ]
 
 
 
 
 
 
 
86
  }
87
+ ]
 
 
88
  )
89
+ print(output["choices"][0]["message"]["content"])
90
  ```
91
 
92
  ### CLI (llama.cpp)
93
 
94
  ```bash
95
+ ./llama-cli -m llava-1.5-7b-hf-q4_k_m.gguf \
96
+ --mmproj mmproj-model-f16.gguf \
97
+ --image photo.jpg \
98
+ -p "Describe this image in detail." -n 256
 
 
 
 
 
99
  ```
100
 
101
+ ### Ollama
102
 
103
  ```bash
104
+ # Create a Modelfile
105
+ FROM ./llava-1.5-7b-hf-q4_k_m.gguf
106
+ TEMPLATE "{{ .Prompt }}"
107
 
 
108
  ollama create sakthai-vision -f Modelfile
109
+ ollama run sakthai-vision "Describe this image." --image photo.jpg
 
 
 
 
110
  ```
111
 
 
 
 
 
 
112
  ---
113
 
114
+ ## Architecture
115
 
116
+ | Component | Detail |
117
+ |-----------|--------|
118
+ | **Model** | LLaVA-1.5-7B (Vicuna-7B-v1.5 LLM + CLIP ViT-L/14 vision encoder) |
119
+ | **Vision encoder** | CLIP ViT-L/14, 336×336 input, 1024 embedding dim, 24 layers |
120
+ | **Language model** | LLaMA-based, ~7B parameters, 32 layers, 4096 hidden |
121
+ | **Projector** | 2-layer MLP (CLIP → LLM token space) |
122
+ | **Original size** | ~6.7 GB (fp16) |
123
+ | **Quantized size** | ~4 GB (GGUF Q4_K_M) + ~400 MB (mmproj) |
124
+ | **Context window** | 4,096 tokens |
125
+ | **Hardware** | 8 GB+ RAM recommended |
126
 
127
+ ---
 
 
 
 
 
 
 
 
128
 
129
+ ## Evaluation (Upstream LLaVA-1.5-7B)
130
 
131
+ These are published upstream benchmarks. As a Q4_K_M quantized repackaging, results should fall within ~1-2%.
132
 
133
+ | Task | Dataset | Score |
134
+ |:-----|:--------:|:-----:|
135
+ | Visual QA | VQAv2 | 78.5% |
136
+ | Captioning | COCO Captions | CIDEr 110.1 |
137
+ | Visual Reasoning | GQA | 62.0% |
138
+ | Text-oriented VQA | TextVQA | 58.2% |
139
+ | Science diagrams | ScienceQA | 89.5% |
140
+ | Visual Perception | MMBench | 66.2% |
141
 
142
+ **Own benchmarks are pending.** This is a repackaging/quantization of upstream LLaVA, not a retrained model.
 
 
 
 
 
 
 
143
 
144
  ---
145
 
146
+ ## Pipeline Integration
147
 
148
+ | Stage | Model | Role |
149
+ |-------|-------|------|
150
+ | 👁️ **See** | **SakThai Vision 7B** ⬅ | **Image→text, visual QA** |
151
+ | 🔍 Retrieve | [Embedding Multilingual](https://huggingface.co/Nanthasit/sakthai-embedding-multilingual) | Cross-lingual search |
152
+ | 🧠 Reason | [Context 1.5B](https://huggingface.co/Nanthasit/sakthai-context-1.5b-merged) or [7B](https://huggingface.co/Nanthasit/sakthai-context-7b-merged) | Tool-calling & reasoning |
153
+ | 🎤 Speak | [TTS Model](https://huggingface.co/Nanthasit/sakthai-tts-model) | Text-to-speech |
154
 
155
  ---
156
 
157
+ ## Files
158
+
159
+ | File | Size | Purpose |
160
+ |:-----|:----:|:--------|
161
+ | `llava-1.5-7b-hf-q4_k_m.gguf` | ~4 GB | Quantized language + vision weights |
162
+ | `mmproj-model-f16.gguf` | ~400 MB | CLIP vision projector |
163
 
164
+ ---
 
 
 
 
 
 
 
 
 
 
165
 
166
+ ## The House of Sak 🏠
167
 
168
+ Until this model, every SakThai agent was blind — it could reason, use tools, and answer questions but couldn't look at screenshots, read whiteboards, or recognize faces. This model was built on free Colab GPUs from a shelter in Cork, Ireland, with no budget and no guarantee of success. LLaVA-1.5-7B's vision encoder was merged, quantized to Q4_K_M GGUF for CPU use, and packed to about 4 GB.
169
 
170
+ > *"We are one family — and becoming more."* — Beer (beer-sakthai)
 
 
 
 
171
 
172
  ---
173
 
174
+ ## Support
175
 
176
+ - ⭐ Leave a like — the first user like on a SakThai model was on this model ❤️
177
+ - 🐛 Report issues on [GitHub](https://github.com/beer-sakthai/Sak-Family-Agent)
178
+ - 🔄 Share with anyone building privacy-focused multimodal apps
179
 
180
  ---
181
 
182
+ ## License
 
 
183
 
184
+ Based on LLaVA-1.5-7B (LLaMA 2 Community License) + CLIP vision encoder. GGUF via llama.cpp. Review upstream licenses before commercial use.
 
 
 
185
 
186
+ ---
 
 
 
 
187
 
188
+ *Built from a shelter in Cork, Ireland. Every download supports a family building AI for the edge.*