kaz-embedder: universal Kazakh text embeddings to RAG classic

kaz-embedder is a Kazakh embedding model for retrieval, semantic similarity, NLI/paraphrase detection, classification and clustering. It is one SentenceTransformer with a Router over three towers:

route tower how to call
query BAAI/bge-m3 with a fine-tuned query tower (distilled from the cross-encoder BAAI/bge-reranker-v2-m3; weight soup of 6 runs) model.encode_query(...)
document BAAI/bge-m3, unchanged (existing bge-m3 document indexes stay valid) model.encode_document(...)
sym (default) intfloat/multilingual-e5-large-instruct, fine-tuned for Kazakh similarity / NLI / paraphrase (weight soup of 4 runs) model.encode(texts, prompt_name=...)

Requires sentence-transformers>=6.1.

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BekbolatShanakbay/kaz-embedder")

# Retrieval: queries and documents go through different towers
q = model.encode_query(["Жыл сайынғы еңбек демалысы неше күн?"])
d = model.encode_document([
    "Жыл сайынғы ақы төленетін еңбек демалысы күнтізбелік жиырма төрт күн болып белгіленеді.",
    "Салық кодексіне сәйкес көлік салығы жыл сайын төленеді.",
])
print(model.similarity(q, d))

# Symmetric tasks: pick the task prompt
e = model.encode(["Ер адам гитарада ойнап отыр.", "Бір кісі гитара тартып отыр."], prompt_name="sts")
print(model.similarity(e[:1], e[1:]))
# prompt_name: "sts" | "pair" | "nli" | "classification" | "clustering" | "bitext"

Evaluation: KazMTEB-v0 test split

Full leaderboard with all Kazakh embedding models on the Hub, the protocol and per-task scores: BekbolatShanakbay/kazmteb-leaderboard. Every model was evaluated under identical conditions: the same 21 test tasks, one run per model, the prompts its card prescribes, and inputs up to min(2048, the model's positional capacity) tokens. All model choices for this model were made on the dev split, and test was run once.

KK is the mean over the Kazakh families R (retrieval), S (STS), P (pair classification), C (classification) and CL (clustering). X is the cross-lingual family. Δ KK is a paired bootstrap of this model minus the row's model (10,000 resamples, 95% CI).

model KK R S P C CL X Δ KK [95% CI]
kaz-embedder (this model) 72.99 67.52 80.91 88.26 73.72 54.56 92.31 —
intfloat/multilingual-e5-large-instruct 70.49 63.98 77.41 85.23 69.68 56.16 90.31 +2.50 [+2.05, +2.81]
BAAI/bge-m3 67.90 67.26 73.12 85.23 71.66 42.26 90.80 +5.09 [+4.50, +5.44]
alphaedge-ai/bge-m3-kaz-32768 67.70 67.24 73.15 85.21 71.64 41.27 90.70 +5.29 [+4.65, +5.60]
Darmm/darmm-embedding-multilingual 65.21 64.47 72.77 82.42 68.62 37.78 89.22 +7.78 [+7.41, +8.43]
shyngys879/kazakh-e5-rag-embedding ⚠️ in-domain 64.89 55.70 70.72 76.54 71.55 49.92 82.85 +8.11 [+7.75, +8.90]
sultanbi/e5-base-kazakh 63.90 55.39 70.19 75.53 71.94 46.46 84.71 +9.09 [+8.31, +9.38]
Nurlykhan/kazembed-v5 ⚠️ in-domain 63.74 52.18 70.21 75.58 71.23 49.49 79.54 +9.26 [+8.65, +9.86]
Tim2190/granite-278m-kk ⚠️ in-domain 63.15 53.88 70.90 75.53 69.13 46.29 79.48 +9.85 [+9.60, +10.81]
intfloat/multilingual-e5-base 60.35 50.88 70.93 82.94 61.42 35.58 77.81 +12.64 [+12.06, +13.29]

⚠️ in-domain: the model card lists KazQAD (the source of task R1) among the training data.

The leaderboard also lists multilingual-e5-small, Darmm/darmm-embed-kazakh-v2, crossroderick/minidalalm and Torekhan/sentence_similarity_model, all below 62 KK.

Strengths and limitations:

  • Kazakh aggregate: best of all evaluated models, and the margin over every model is statistically significant.
  • STS, pair classification, classification: significantly better than every other model.
  • Retrieval (R): best of all models, but only +0.3 over bge-m3, which it is built on. That margin is within noise.
  • Clustering (CL): multilingual-e5-large-instruct is 1.6 points higher, mostly on SIB-200 topic clustering.
  • Home advantage: the leaderboard authors built this model, so the results carry a home-team caveat, disclosed on the leaderboard page.

Training data

The model is zero-shot with respect to KazMTEB:

  • Excluded from training: evaluation data and exam question banks, namely KazQAD questions and passages, ЕНТ/UNT, KazMMLU, WebFAQ and STS-B.
  • Leak checks on every training text:
    • a registry of all evaluation items, matched exactly, by 6-grams and by word Jaccard;
    • a semantic filter that drops texts with bge-m3 cosine ≥ 0.88 to any evaluation query;
    • an exact check of the English originals of translated NLI data.

Data used:

  • Retrieval (query tower). Real Kazakh texts: Wikipedia, legal acts and codes, news, government FAQ and cultural Q&A. Queries were:

    • questions written about these passages by Claude (Anthropic);
    • real questions;
    • structural pairs (title → lead, article title → text);
    • sentences taken from the passages themselves.

    The teacher cross-encoder BAAI/bge-reranker-v2-m3 labelled the candidates.

  • Symmetric tower. Kazakh NLI, paraphrase and adversarial-paraphrase data (translated SNLI, QQP and PAWS), plus paraphrase/entailment/contradiction pairs written by Claude about real sentences.

License and attribution

The model weights are released under Apache-2.0. The base models are BAAI/bge-m3 (MIT) and intfloat/multilingual-e5-large-instruct (MIT). Their copyright notices are kept, and use of this model is also subject to their terms.

Notes

  • Size: three XLM-RoBERTa-large towers, about 3 × 560M parameters (3.4 GB in fp16).
  • Scope: Kazakh was the only target. Other languages were not optimised.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BekbolatShanakbay/kaz-embedder

Base model

BAAI/bge-m3
Finetuned
(573)
this model