How to use from
Docker Model Runner
docker model run hf.co/IFM/K2-Horizon-0.9B
Quick Links

K2-Horizon-0.9B

K2-Horizon-0.9B is the compact dense member of the K2-Horizon family: a 0.9B-class decoder-only model with a 128K context window.

K2-Horizon-0.9B benchmark results

K2-Horizon-0.9B Highlights

  • Compact reasoning model. A 0.9B-class dense model evaluated across mathematics, coding, science, and tool-use benchmarks.
  • 128K context. Supports up to 131,072 tokens with YaRN RoPE scaling.
  • Multi-teacher distillation. Trained with domain teachers for math and code, STEM, and instruction following.
  • Fully open. Training data/recipe and the training code will be made public.

Benchmark Results

The chart at the top of this card shows K2-Horizon-0.9B against selected reference models. The table below lists every comparison model used in the figure.

Full Results

Reference models
K2-Horizon-0.9BQwen3.5-0.8BOpenBMB-1BQwen3.5-2B
# Params0.9B0.8B1B2B
# Activated params0.9B0.8B1B2B
ArchitectureDenseDenseDenseDense
Math
AIME 2025
Competition mathematics
41.71.040.434.2
AIME 2026
Competition mathematics
48.50.240.438.8
HMMT Feb 2026
Competition mathematics
25.80.623.322.7
Scientific Reasoning
GPQA Diamond
Graduate-level science QA
27.311.926.354.9
Coding
HumanEval+
Code generation
79.916.565.275.6
MBPP+
Code generation
68.035.460.667.7
LiveCodeBench v6
Competitive coding
37.46.633.529.8
Agents
BFCL v4
Function calling
28.025.325.243.6

Scores in %. Bold highlights K2-Horizon-0.9B; Qwen3.5-2B is included as a larger reference model. Protocol and provenance details are in the Technical Appendix.

Quickstart

Serving

vLLM (source at PR #53806, commit d9fd5f11):

vllm serve IFM/K2-Horizon-0.9B \
  --trust-remote-code \
  --dtype bfloat16 \
  --max-model-len 131072 \
  --hf-overrides '{"rope_parameters":{rope_type: yarn, factor: 16, original_max_position_embeddings: 8192, rope_theta: 1000000, beta_fast: 128, beta_slow: 4}' \
  --gpu-memory-utilization 0.85 \
  --tensor-parallel-size 1 \
  --reasoning-parser k2_horizon \
  --enable-auto-tool-choice \
  --tool-call-parser k2_horizon

Use an exact branch name from the inventory with vLLM's --revision option. For example, --revision pretrain_600000 selects the final checkpoint of Pretraining, at step 600,000.

SGLang, from a source checkout that includes sgl-project/sglang#37654. This is the recipe validated in the SGLang K2 Horizon cookbook:

sglang serve \
  --model-path IFM/K2-Horizon-0.9B \
  --revision 9b9ec1f7e17f62ed218df542687a144116219d84 \
  --tp 1 \
  --dtype bfloat16 \
  --attention-backend fa3 \
  --reasoning-parser k2_horizon \
  --host 0.0.0.0 \
  --port 30000

API Usage

Recommended settings: reasoning_effort="high", temperature=0.6, top_p=0.95, and at least 32,768 output tokens. Reasoning depth is selected per request through chat_template_kwargs. Thinking is returned in reasoning_content and the answer in content.

from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="IFM/K2-Horizon-0.9B",
    messages=[{"role": "user", "content": "Explain the result step by step."}],
    temperature=0.6,
    top_p=0.95,
    max_tokens=32768,
    extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)

Transformers

Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "IFM/K2-Horizon-0.9B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)

inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training Overview

The table below lists the training stages in order and the purpose of each stage.

Training steps are counted within each stage or phase. Token budgets cover only the additional training in that stage or phase.

Each stage or phase continues from the final checkpoint of the preceding stage or phase.

During RL, training branches into seven expert models, which are then merged, as described below.

Training stage Training steps Training tokens Sequence length Purpose
Pretraining 600000 5T 8K Pretraining.
Midtraining — Stage 1 75000 625B 32K Context extension.
Midtraining — Stage 2 47684 397B 128K Context extension.
RL To be updated To be updated 128K We trained seven expert models from the final checkpoint of Midtraining Stage 2: math1, code1, math2a, math2b, code2, IF, and stem. We then merged the expert models.
MOPD 249 To be updated 128K Resolve structural interference and performance degradation caused by weight merging, aligning multi-domain specialist capabilities in the behavioral space via on-policy distillation.

Release Artifacts

The tables below list the release artifacts for K2-Horizon-0.9B, their availability, and the expected release dates for remaining items.

Last updated: 2026-09-11

Status:

  • Available — fully released for the scope listed;
  • Partial — some items are available, with remaining items listed in the notes;
  • In Progress — being prepared for release but not yet available.

Artifact Index

Artifact Link Status Remaining items / expected availability
Model card Hugging Face Available N/A
Training logs W&B Available N/A
Blog post Blog post Available N/A
Checkpoints Checkpoint inventory Partial See details below
Technical report Not yet available In Progress End of September 2026
Code repository GitHub In Progress End of September 2026

Checkpoint Inventory

Model repository: IFM/K2-Horizon-0.9B

Branch names below refer to this repository. Patterns containing * group branches by training stage or phase. The * is a placeholder for a training-step number, not a literal branch name. Intermediate checkpoint groups exclude the final checkpoint listed separately; a pattern does not imply that a checkpoint is available at every step.

For example, pretrain_600000 is the checkpoint saved at training step 600,000 within Pretraining stage, and is the final checkpoint of that stage. The numeric suffix is the step within the named stage, not the cumulative step across all training. Thus, mid_1_75000 refers to step 75,000 within Midtraining Stage 1.

For a partially released group, the available checkpoints and the remaining checkpoints are listed in the notes.

Checkpoint Branch / repository Status Remaining items / expected availability
Pretrain Intermediate Checkpoints pretrain_* Available N/A
Pretrain Final Checkpoint pretrain_600000 Available N/A
Midtrain Stage 1 Intermediate Checkpoints mid_1_* Available N/A
Midtrain Stage 1 Final Checkpoint mid_1_75000 Available N/A
Midtrain Stage 2 Intermediate Checkpoints mid_2_* Available N/A
Midtrain Stage 2 Final Checkpoint mid_2_47684 Available N/A
RL Math1 Expert Checkpoint rl_math1 In Progress Mid-September 2026
RL Code1 Expert Checkpoint rl_code1 In Progress Mid-September 2026
RL Math2a Expert Checkpoint rl_math2a In Progress Mid-September 2026
RL Math2b Expert Checkpoint rl_math2b In Progress Mid-September 2026
RL Code2 Expert Checkpoint rl_code2 In Progress Mid-September 2026
RL IF Expert Checkpoint rl_if In Progress Mid-September 2026
RL Stem Expert Checkpoint rl_stem In Progress Mid-September 2026
RL Merged Final Checkpoint rl_merged Available N/A
RL MOPD Final Checkpoint rl_mopd Available N/A

Best Practices

  1. Reasoning effort: always high. All reported results use high reasoning effort. Pass {"chat_template_kwargs": {"reasoning_effort": "high"}} on every request; medium and low trade accuracy for speed and are not recommended for evaluation.
  2. Sampling parameters. temperature=0.6, top_p=0.95.
  3. Output length. Allow at least 32,768 output tokens so reasoning is never cut off. Truncated reasoning is a failed response, not a shorter one.
  4. Serving. Use the validated SGLang recipe above: BF16, TP=1, FlashAttention-3. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the SGLang cookbook.
  5. Parsers. Enable the k2_horizon reasoning parser for chat, and add the k2_horizon tool-call parser for agent use. Leave both off for plain completion-style generation.
  6. Revisions. main is the MOPD release checkpoint; mid1_75k and mid2_47k preserve the context-extension stages.

Citation

@misc{k2horizon2026,
  title  = {Introducing K2 Horizon: Frontier Performance, Radically Open},
  author = {{IFM Team}},
  year   = {2026},
  url    = {https://ifm.ai/blog/k2/},
}
Downloads last month
13,938
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for IFM/K2-Horizon-0.9B

Adapters
2 models
Finetunes
2 models
Quantizations
7 models

Space using IFM/K2-Horizon-0.9B 1

Collection including IFM/K2-Horizon-0.9B