Instructions to use IFM/K2-Horizon-0.9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IFM/K2-Horizon-0.9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IFM/K2-Horizon-0.9B", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("IFM/K2-Horizon-0.9B", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IFM/K2-Horizon-0.9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IFM/K2-Horizon-0.9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-0.9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/IFM/K2-Horizon-0.9B
- SGLang
How to use IFM/K2-Horizon-0.9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IFM/K2-Horizon-0.9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-0.9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IFM/K2-Horizon-0.9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IFM/K2-Horizon-0.9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use IFM/K2-Horizon-0.9B with Docker Model Runner:
docker model run hf.co/IFM/K2-Horizon-0.9B
K2-Horizon-0.9B
K2-Horizon-0.9B is the compact dense member of the K2-Horizon family: a 0.9B-class decoder-only model with a 128K context window.
K2-Horizon-0.9B Highlights
- Compact reasoning model. A 0.9B-class dense model evaluated across mathematics, coding, science, and tool-use benchmarks.
- 128K context. Supports up to 131,072 tokens with YaRN RoPE scaling.
- Multi-teacher distillation. Trained with domain teachers for math and code, STEM, and instruction following.
- Fully open. Training data/recipe and the training code will be made public.
Benchmark Results
The chart at the top of this card shows K2-Horizon-0.9B against selected reference models. The table below lists every comparison model used in the figure.
Full Results
| Reference models | ||||
|---|---|---|---|---|
| K2-Horizon-0.9B | Qwen3.5-0.8B | OpenBMB-1B | Qwen3.5-2B | |
| # Params | 0.9B | 0.8B | 1B | 2B |
| # Activated params | 0.9B | 0.8B | 1B | 2B |
| Architecture | Dense | Dense | Dense | Dense |
| Math | ||||
AIME 2025 Competition mathematics | 41.7 | 1.0 | 40.4 | 34.2 |
AIME 2026 Competition mathematics | 48.5 | 0.2 | 40.4 | 38.8 |
HMMT Feb 2026 Competition mathematics | 25.8 | 0.6 | 23.3 | 22.7 |
| Scientific Reasoning | ||||
GPQA Diamond Graduate-level science QA | 27.3 | 11.9 | 26.3 | 54.9 |
| Coding | ||||
HumanEval+ Code generation | 79.9 | 16.5 | 65.2 | 75.6 |
MBPP+ Code generation | 68.0 | 35.4 | 60.6 | 67.7 |
LiveCodeBench v6 Competitive coding | 37.4 | 6.6 | 33.5 | 29.8 |
| Agents | ||||
BFCL v4 Function calling | 28.0 | 25.3 | 25.2 | 43.6 |
Scores in %. Bold highlights K2-Horizon-0.9B; Qwen3.5-2B is included as a larger reference model. Protocol and provenance details are in the Technical Appendix.
Quickstart
Serving
vLLM (source at PR #53806, commit d9fd5f11):
vllm serve IFM/K2-Horizon-0.9B \
--trust-remote-code \
--dtype bfloat16 \
--max-model-len 131072 \
--hf-overrides '{"rope_parameters":{rope_type: yarn, factor: 16, original_max_position_embeddings: 8192, rope_theta: 1000000, beta_fast: 128, beta_slow: 4}' \
--gpu-memory-utilization 0.85 \
--tensor-parallel-size 1 \
--reasoning-parser k2_horizon \
--enable-auto-tool-choice \
--tool-call-parser k2_horizon
Use an exact branch name from the inventory with vLLM's --revision option. For example, --revision pretrain_600000 selects the final checkpoint of Pretraining, at step 600,000.
SGLang, from a source checkout that includes sgl-project/sglang#37654. This is the recipe validated in the SGLang K2 Horizon cookbook:
sglang serve \
--model-path IFM/K2-Horizon-0.9B \
--revision 9b9ec1f7e17f62ed218df542687a144116219d84 \
--tp 1 \
--dtype bfloat16 \
--attention-backend fa3 \
--reasoning-parser k2_horizon \
--host 0.0.0.0 \
--port 30000
API Usage
Recommended settings:
reasoning_effort="high",temperature=0.6,top_p=0.95, and at least 32,768 output tokens. Reasoning depth is selected per request throughchat_template_kwargs. Thinking is returned inreasoning_contentand the answer incontent.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
response = client.chat.completions.create(
model="IFM/K2-Horizon-0.9B",
messages=[{"role": "user", "content": "Explain the result step by step."}],
temperature=0.6,
top_p=0.95,
max_tokens=32768,
extra_body={"chat_template_kwargs": {"reasoning_effort": "high"}},
)
message = response.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print("Answer:", message.content)
Transformers
Validated with Transformers 5.15.0, PyTorch 2.13.0, Safetensors 0.8.0.
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "IFM/K2-Horizon-0.9B"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id, device_map="auto", dtype="bfloat16", low_cpu_mem_usage=True, trust_remote_code=True
)
inputs = tokenizer("Explain why long-context evaluation is difficult.", return_tensors="pt").to(model.device)
inputs.pop("token_type_ids", None)
outputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Training Overview
The table below lists the training stages in order and the purpose of each stage.
Training steps are counted within each stage or phase. Token budgets cover only the additional training in that stage or phase.
Each stage or phase continues from the final checkpoint of the preceding stage or phase.
During RL, training branches into seven expert models, which are then merged, as described below.
| Training stage | Training steps | Training tokens | Sequence length | Purpose |
|---|---|---|---|---|
| Pretraining | 600000 | 5T | 8K | Pretraining. |
| Midtraining — Stage 1 | 75000 | 625B | 32K | Context extension. |
| Midtraining — Stage 2 | 47684 | 397B | 128K | Context extension. |
| RL | To be updated | To be updated | 128K | We trained seven expert models from the final checkpoint of Midtraining Stage 2: math1, code1, math2a, math2b, code2, IF, and stem. We then merged the expert models. |
| MOPD | 249 | To be updated | 128K | Resolve structural interference and performance degradation caused by weight merging, aligning multi-domain specialist capabilities in the behavioral space via on-policy distillation. |
Release Artifacts
The tables below list the release artifacts for K2-Horizon-0.9B, their availability, and the expected release dates for remaining items.
Last updated: 2026-09-11
Status:
- Available — fully released for the scope listed;
- Partial — some items are available, with remaining items listed in the notes;
- In Progress — being prepared for release but not yet available.
Artifact Index
| Artifact | Link | Status | Remaining items / expected availability |
|---|---|---|---|
| Model card | Hugging Face | Available | N/A |
| Training logs | W&B | Available | N/A |
| Blog post | Blog post | Available | N/A |
| Checkpoints | Checkpoint inventory | Partial | See details below |
| Technical report | Not yet available | In Progress | End of September 2026 |
| Code repository | GitHub | In Progress | End of September 2026 |
Checkpoint Inventory
Model repository: IFM/K2-Horizon-0.9B
Branch names below refer to this repository. Patterns containing * group branches by training stage or phase. The * is a placeholder for a training-step number, not a literal branch name. Intermediate checkpoint groups exclude the final checkpoint listed separately; a pattern does not imply that a checkpoint is available at every step.
For example, pretrain_600000 is the checkpoint saved at training step 600,000 within Pretraining stage, and is the final checkpoint of that stage. The numeric suffix is the step within the named stage, not the cumulative step across all training. Thus, mid_1_75000 refers to step 75,000 within Midtraining Stage 1.
For a partially released group, the available checkpoints and the remaining checkpoints are listed in the notes.
| Checkpoint | Branch / repository | Status | Remaining items / expected availability |
|---|---|---|---|
| Pretrain Intermediate Checkpoints | pretrain_* |
Available | N/A |
| Pretrain Final Checkpoint | pretrain_600000 |
Available | N/A |
| Midtrain Stage 1 Intermediate Checkpoints | mid_1_* |
Available | N/A |
| Midtrain Stage 1 Final Checkpoint | mid_1_75000 |
Available | N/A |
| Midtrain Stage 2 Intermediate Checkpoints | mid_2_* |
Available | N/A |
| Midtrain Stage 2 Final Checkpoint | mid_2_47684 |
Available | N/A |
| RL Math1 Expert Checkpoint | rl_math1 |
In Progress | Mid-September 2026 |
| RL Code1 Expert Checkpoint | rl_code1 |
In Progress | Mid-September 2026 |
| RL Math2a Expert Checkpoint | rl_math2a |
In Progress | Mid-September 2026 |
| RL Math2b Expert Checkpoint | rl_math2b |
In Progress | Mid-September 2026 |
| RL Code2 Expert Checkpoint | rl_code2 |
In Progress | Mid-September 2026 |
| RL IF Expert Checkpoint | rl_if |
In Progress | Mid-September 2026 |
| RL Stem Expert Checkpoint | rl_stem |
In Progress | Mid-September 2026 |
| RL Merged Final Checkpoint | rl_merged |
Available | N/A |
| RL MOPD Final Checkpoint | rl_mopd |
Available | N/A |
Best Practices
- Reasoning effort: always
high. All reported results use high reasoning effort. Pass{"chat_template_kwargs": {"reasoning_effort": "high"}}on every request;mediumandlowtrade accuracy for speed and are not recommended for evaluation. - Sampling parameters.
temperature=0.6,top_p=0.95. - Output length. Allow at least 32,768 output tokens so reasoning is never cut off. Truncated reasoning is a failed response, not a shorter one.
- Serving. Use the validated SGLang recipe above: BF16, TP=1, FlashAttention-3. Full recipes for every K2-Horizon size, with measured H200 latency and throughput, are in the SGLang cookbook.
- Parsers. Enable the
k2_horizonreasoning parser for chat, and add thek2_horizontool-call parser for agent use. Leave both off for plain completion-style generation. - Revisions.
mainis the MOPD release checkpoint;mid1_75kandmid2_47kpreserve the context-extension stages.
Citation
@misc{k2horizon2026,
title = {Introducing K2 Horizon: Frontier Performance, Radically Open},
author = {{IFM Team}},
year = {2026},
url = {https://ifm.ai/blog/k2/},
}
- Downloads last month
- 13,938
docker model run hf.co/IFM/K2-Horizon-0.9B