Yigdzin-1 β€” GGUF (llama.cpp)

GGUF conversion of BDRC/tibetan-ocr ("Yigdzin-1"), a Tibetan OCR vision-language model fine-tuned from PaddlePaddle/PaddleOCR-VL-1.6, for use with llama.cpp / llama-server (mtmd multimodal support). Used by the local-ocr desktop app.

Files

  • yigdzin1.gguf β€” the language model (Q8_0-equivalent precision inherited from the source checkpoint's conversion; same conversion path as PaddlePaddle/PaddleOCR-VL-1.6-GGUF)
  • yigdzin1-mmproj.gguf β€” the vision projector (--mmproj)

Conversion notes

Converted with a stock (unmodified) llama.cpp build (b10776) using its conversion/ernie.py PaddleOCR-VL converter, with one compatibility fix: this checkpoint ships processor_config.json with size.shortest_edge/size.longest_edge instead of the flat min_pixels/max_pixels keys the base PaddleOCR-VL-1.6 uses, which the stock converter doesn't handle. No change was made to inference-time behavior β€” the produced GGUF carries no non-standard metadata and loads/runs identically to any other PaddleOCR-VL-arch GGUF on stock llama-server.

(An earlier, separate experiment explored whether Yigdzin-1's image-token positional encoding needs a non-standard "sequential" mode at inference time β€” relevant for whole-page images. It was found unnecessary for the line-level crops this conversion targets; see the source app's repository for details.)

Usage

llama-server -m yigdzin1.gguf --mmproj yigdzin1-mmproj.gguf

Then send a chat-completion request with an image (line crop) and the prompt "Extract all Tibetan text.".

License

Apache-2.0, inherited from BDRC/tibetan-ocr.

Downloads last month
91
GGUF
Model size
0.4B params
Architecture
paddleocr
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for nakamura196/yigdzin1-gguf

Finetuned
BDRC/tibetan-ocr
Quantized
(1)
this model