--- license: apache-2.0 language: - en - zh - ko - ja - multilingual library_name: transformers pipeline_tag: text-generation tags: - darwin - darwin-v7 - evolutionary-merge - merge - mergekit - reasoning - advanced-reasoning - chain-of-thought - thinking - qwen3.6 - qwen - claude-opus - distillation - multilingual - gpqa - benchmark - open-source - apache-2.0 - hybrid-vigor - proto-agi - vidraft - eval-results base_model_relation: merge model-index: - name: Darwin-28B-Opus results: - task: type: question-answering name: Question Answering dataset: name: GPQA Diamond type: Idavidrein/gpqa config: gpqa_diamond metrics: - type: accuracy value: 88.89 name: Accuracy --- > ### ๐ฑ Run it on your phone or a GPU-less PC โ **POCKET** ยท ๐ **[Try it live (CPU chat)](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU)** > VIDRAFT's on-device family: a 35B model that runs on **iPhone** and on **CPU with no GPU** โ stock `llama.cpp`, no fork. > > [](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) [](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6) [](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) [](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) [](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) > # Darwin-28B-Opus โ Qwen3.6-27B ร Opus-Distilled Evolutionary Merge
> Qwen3.6-27B dense ยท 27.6B parameters ยท Hybrid Linear/Full Attention ยท BF16 ยท Thinking Mode ยท Apache 2.0 > **Darwin V7 evolutionary merge: Father ร Opus-distilled Mother โ 88.89% on GPQA Diamond (3-stage adaptive evaluation)** --- ## Abstract **Darwin-28B-Opus** is the first reasoning model of the Darwin series built on the **Qwen3.6 generation** backbone. Produced by the Darwin V7 evolutionary breeding engine from two publicly available parents, it combines the strong bilingual reasoning of Qwen3.6-27B with Claude Opus 4-style chain-of-thought distilled behaviour. On the **GPQA Diamond** graduate-level reasoning benchmark (198 PhD-level questions), Darwin-28B-Opus scores **88.89 %** under the standard 3-stage adaptive evaluation, slightly edging out its larger MoE sibling Darwin-36B-Opus (88.4 %) and clearly surpassing its Qwen3.5-generation counterpart Darwin-27B-Opus (86.9 %). --- ## ๐งฌ Model Lineage | Role | Model | Role in the Merge | |:---:|:---|:---| | **Father (็ถ)** | [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) | Qwen3.6 generation dense backbone with hybrid linear/full attention. | | **Mother (ๆฏ)** | [`rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled`](https://huggingface.co/rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled) | Claude Opus reasoning-distilled variant of the same backbone (Jackrong-style distillation, 14 k traces). | | **Offspring** | **`Darwin-28B-Opus`** (this model) | Darwin V7 evolutionary merge; Qwen3.6 architecture retained, Opus reasoning style inherited. | > **Why 28B?** The `28B` label denotes the Qwen3.6-generation member of the Darwin lineup (`+1` over the Qwen3.5-era `Darwin-27B-Opus`). > The actual parameter count is **27.6 B**, and the architecture exactly follows Qwen3.6-27B. --- ## โ๏ธ Technical Specifications | Component | Value | |:---|:---| | Architecture | `Qwen3_5ForConditionalGeneration` (Qwen3.6 generation, hybrid linear + full attention) | | Parameters | **27.6 B** (BF16) | | Hidden size | 5 120 | | Intermediate size | 17 408 | | Head dim | 256 | | Layers | 64 (3 linear : 1 full attention, `full_attention_interval = 4`) | | Precision | bfloat16 | | Context length | Inherited from base (long-chain reasoning supported) | | License | Apache 2.0 | --- ## ๐ Benchmark โ GPQA Diamond (198 questions) Darwin-28B-Opus is evaluated under our standard **3-stage adaptive evaluation** protocol, identical to the protocol used across the Darwin series. | Stage | Decoding Protocol | Cost | **Accuracy** | |:---:|:---|:---:|:---:| | **Stage 1** | Single-shot greedy baseline | 1ร | **74.75 %** (148 / 198) | | **Stage 2** | Majority vote ร8 at temperature 0.7 on Stage-1 wrongs | 8ร | **83.84 %** (166 / 198) | | **Stage 3** | Adaptive ensemble refinement (close-tie tiebreaker + iterative MTI on residual hard questions) | โ 20ร | **๐ฅ 88.89 %** (176 / 198) | **Key performance indicators**: - Stage 1 โ Stage 3: **+14.14 %p** through adaptive protocol - vs Darwin-27B-Opus (86.9 %): **+1.99 %p** - vs Darwin-36B-Opus (88.4 %): **+0.49 %p** - vs Darwin-31B-Opus (85.9 %): **+2.99 %p** --- ## ๐ Usage ### Standard inference (Stage 1 baseline) ```python from transformers import AutoTokenizer, AutoModelForCausalLM import torch tok = AutoTokenizer.from_pretrained( "FINAL-Bench/Darwin-28B-Opus", trust_remote_code=True, ) model = AutoModelForCausalLM.from_pretrained( "FINAL-Bench/Darwin-28B-Opus", torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True, ) messages = [ {"role": "user", "content": "Solve: If f(x) = xยณ โ 3x + 2, find all critical points and classify them."} ] text = tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) inputs = tok(text, return_tensors="pt").to(model.device) outputs = model.generate(**inputs, max_new_tokens=2048, do_sample=False) print(tok.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)) ``` ### Enhanced accuracy (Stage 2-3 adaptive) For leaderboard-grade accuracy, combine: 1. Stage 1 greedy baseline, 2. Stage 2 maj@8 temperature sampling on low-confidence answers, 3. Stage 3 adaptive refinement on still-disputed answers. Reference implementation is provided in the Darwin-series evaluation harness. --- ## ๐ฏ Recommended Use-Cases - **Graduate-level STEM reasoning** (GPQA / science qualifying exams) - **Mathematical problem solving** (MATH, AIME-style problems) - **Code generation and debugging** (HumanEval, MBPP) - **Complex multi-step chain-of-thought tasks** - **Bilingual reasoning** (strong English + Korean; also Chinese / Japanese) ## โ ๏ธ Limitations - At 27.6 B parameters in bfloat16, full inference requires โ 55 GB of VRAM (e.g., a single A100-80GB or B200). - Optimised for English first, with secondary support for Korean, Chinese, and Japanese. - Deep Opus-style reasoning traces tend to be verbose โ control with `max_new_tokens` as needed. --- ## ๐ Citation ```bibtex @misc{darwin28b_opus_2026, title = {Darwin-28B-Opus: Evolutionary Merging of Qwen3.6-27B with Claude-Opus-Distilled Reasoning}, author = {FINAL-Bench / Darwin Research Team}, year = {2026}, howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-28B-Opus}}, note = {Darwin V7 ยท Mother-centric Ratio Interpolation merge ยท 88.89 % GPQA Diamond (3-stage)} } ``` --- ## ๐ Related Darwin Models - **Darwin-36B-Opus** โ MoE 36B, Qwen3.6-35B-A3B ร Opus distilled, GPQA 88.4 % - **Darwin-31B-Opus** โ 31B dense, multilingual-strong reasoning, GPQA 85.9 % - **Darwin-27B-Opus** โ 27B dense (Qwen3.5 generation), GPQA 86.9 % - **Darwin-9B-NEG** โ 9B with Native Entropy Gating, GPQA 84.3 % - **Darwin-9B-Opus** โ the Qwen3.5-9B Darwin member - **Darwin-4B-Genesis** โ smallest Darwin member --- ## ๐ฑ Related โ POCKET (run a 35B model on-device) Want Darwin-class reasoning without a datacenter GPU? **POCKET** is VIDRAFT's on-device family โ a **35B model that runs on a phone and on a GPU-less PC** using stock `llama.cpp` (no fork, no CUDA, no cloud). On a free CPU it generates **~3.4ร faster than Bonsai**, the most-downloaded on-device model (2M+), at matched quality (HellaSwag 61.0 % vs 60.0 %, a statistical tie). - โถ๏ธ **Live CPU chat (try it now):** https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU - ๐ **POCKET collection:** https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6 - ๐ฆ **POCKET-35B-GGUF:** https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF - ๐ฐ๐ท **POCKET-KR-MLX** (iPhone / Mac): https://huggingface.co/FINAL-Bench/POCKET-KR-MLX - ๐ฌ๐ง **POCKET-EN-GGUF:** https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF --- This model is introduced in [Darwin Family](https://arxiv.org/abs/2605.14386). *Darwin V7 ยท Qwen3.6 generation flagship ยท Sealed 2026-04-25 ยท FINAL-Bench*