AI & ML interests

Darmm (darmm.kz): a personal path into AI for students and developers — math → ML → AI. Pick a concrete goal, a diagnostic maps what you're missing, an AI tutor drills you daily, an ML engineer teaches live weekly. Kazakh, Russian, English. Also: open Kazakh-language models (OCR, embeddings, sentiment).

Recent Activity

R3iwan  updated a Space 11 days ago
Darmm/README
R3iwan  updated a model about 1 month ago
Darmm/darmm-chat-kazakh-8b-lora
R3iwan  updated a collection about 1 month ago
Darmm Text Generation Kazakh
View all activity

Organization Card

Darmm

Darmm is two things, built by one person — R3iwan, an ML/AI engineer and teacher in Astana, Kazakhstan:

  1. darmm.kz — a personal path into AI. For students and developers who want to get into AI and keep hitting the math wall. You pick a concrete goal (a paper, a model, an interview); a 20-minute diagnostic maps what you already know; the program walks the shortest path through math → ML → AI, with an AI tutor that sees your work every day and a live session with an ML engineer every week. In Kazakh, Russian and English. First cohort starts 17 November 2026.
  2. Open models, datasets and benchmarks for underserved languages, starting with Kazakh — everything published under this organization.

Open models: language and document AI for underserved languages

The focus is the neglected part of the AI landscape: languages and domains where global open-source models quietly drop in quality — where tokenizers fragment text, models drift into a dominant neighbor language, and local script, context, and terminology are treated as edge cases.

What I'm working on

  • Text / NLP — LLM fine-tuning, embeddings, classification, and generation for low-resource languages, with a focus on legal and technical terminology. Current work centers on Kazakh and Russian.
  • Tokenization & language adaptation — measuring and closing the gap between how global models handle underserved languages versus how they should: tokenizer efficiency, vocabulary extension, continued pretraining on native-language data. A methodology built on Kazakh, designed to transfer.
  • OCR & document AI — recognition of printed and handwritten text in underrepresented scripts, adapted for legal, educational, and government documents; extraction and validation pipelines on top of it.
  • Retrieval & RAG — embedding models, hybrid search, and RAG/GraphRAG pipelines tuned for non-English corpora, evaluated against curated golden sets.
  • Benchmarks & datasets — open datasets and honest, reproducible evaluation reports of where current models actually stand on languages the leaderboards ignore.

Most work is published as open models, datasets, and benchmarks on Hugging Face. Production-specific components stay closed.

Technical focus

  • LLM adaptation: LoRA/QLoRA fine-tuning, vocabulary extension, domain-specific embeddings, instruction tuning for low-resource settings.
  • OCR: classical and transformer-based OCR pipelines, support for multiple scripts and alphabet variants (currently Kazakh Cyrillic and Latin), document layout and field extraction.
  • Retrieval: embedding benchmarking, hybrid scoring (BM25 + dense), rerankers, RAG and GraphRAG pipelines.
  • Inference: vLLM serving, quantization (AWQ, GGUF), latency optimization for production deployment.
  • Evaluation: CER for OCR, retrieval and generation metrics for RAG (including LLM-as-judge calibrated against human labels), honest reporting of where models fail.

Why start with Kazakh

Kazakh is the native language of over 20 million people, the state language of Kazakhstan, and the language of a growing digital economy. In AI research, it is consistently treated as an afterthought — and it is far from alone.

Tokenizers of major open models fragment Kazakh text far more than English or Russian, degrading both quality and cost. General-purpose LLMs either slip into Russian when prompted in Kazakh, or produce degraded output. OCR systems regularly confuse Kazakh Cyrillic with Russian. Domain terminology — legal, governmental, technical — is where the drop is sharpest.

The same failure pattern repeats across dozens of languages worldwide: a dominant neighbor language the model falls back to, a script the tokenizer wasn't built for, domain vocabulary nobody benchmarked. Kazakh is where Darmm builds and validates the playbook — systematic benchmarking, open models, honest reporting — with the explicit goal of making the methods, tooling, and evaluation practices reusable for other underserved languages.

Coming later

  • Kazakh ↔ Russian / Kazakh ↔ English translation, with the pipeline generalized to other language pairs
  • Multimodal document understanding (vision + language)
  • Extension of the adaptation and benchmarking stack to other Central Asian and low-resource languages
  • Speech (ASR/TTS) — deliberately parked until the text and document stack is solid

Contact

Program and waitlist: darmm.kz. Models and research: github.com/R3iwan · Telegram @rakhinator.