Rewrite model card: BigBang-style structure, expanded benchmark comparison, remove banner
Browse files
README.md
CHANGED
|
@@ -8,50 +8,33 @@ tags:
|
|
| 8 |
pipeline_tag: text-generation
|
| 9 |
---
|
| 10 |
|
| 11 |
-
|
| 12 |
-
<img src="assets/fx-bio-banner.png" alt="Fx-Bio Banner" width="100%"/>
|
| 13 |
-
|
| 14 |
-
<p>
|
| 15 |
-
<a href="https://huggingface.co/endless-frontier/Fx-Bio"><img alt="Model" src="https://img.shields.io/badge/%F0%9F%A4%97%20Model-endless--frontier%2FFx--Bio-536af5"/></a>
|
| 16 |
-
<a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-Apache--2.0-blue.svg"/></a>
|
| 17 |
-
<img alt="Base Model" src="https://img.shields.io/badge/Base-DeepSeek--V4--Flash-6b5ce7"/>
|
| 18 |
-
<img alt="Context" src="https://img.shields.io/badge/Context-1M%20tokens-0d9488"/>
|
| 19 |
-
</p>
|
| 20 |
-
|
| 21 |
-
<p>English ⬇️ | <a href="#chinese">中文 ⬇️</a></p>
|
| 22 |
-
</div>
|
| 23 |
-
|
| 24 |
-
**Fx-Bio-0913** is a biomedical reasoning large language model post-trained on **DeepSeek-V4-Flash**, developed by The Endless Frontier lab. Model weights are openly released under **Apache-2.0**.
|
| 25 |
|
| 26 |
-
##
|
| 27 |
|
| 28 |
-
-
|
| 29 |
-
- 📈 **Consistent gains over the base model** on biomedical reasoning benchmarks, especially on harder problems — e.g. **+5.8** on BioMysteryBench Human-Difficult (avg@5) and **+6.6** on BiominiBench pass@3 over DeepSeek-V4-Flash.
|
| 30 |
-
- 📏 **1M-token context length**, inherited from the base model, for long biomedical documents and multi-hop evidence.
|
| 31 |
-
- 🔓 **Fully open weights** (bf16 safetensors) under Apache-2.0, with Megatron-LM training arguments (`args.json`) included for reproducibility.
|
| 32 |
|
| 33 |
-
##
|
| 34 |
|
| 35 |
-
|
| 36 |
|
| 37 |
-
<
|
| 38 |
-
<img src="assets/fx-bio-eval.png" alt="Fx-Bio
|
| 39 |
-
</
|
| 40 |
-
|
| 41 |
-
**BioMysteryBench**
|
| 42 |
|
| 43 |
-
|
| 44 |
-
|---|---|---|---|
|
| 45 |
-
| Human-Solvable avg@5 | 85.2 | 89.3 | **85.5** |
|
| 46 |
-
| Human-Difficult avg@5 | 31.8 | 45.9 | **37.6** |
|
| 47 |
-
| pass@5 | 82.2 | 86.7 | **86.7** |
|
| 48 |
|
| 49 |
-
|
| 50 |
|
| 51 |
-
|
|
| 52 |
-
|---|---|---|---|
|
| 53 |
-
|
|
| 54 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 55 |
|
| 56 |
## Model Overview
|
| 57 |
|
|
@@ -83,12 +66,12 @@ outputs = model.generate(**inputs, max_new_tokens=1024)
|
|
| 83 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 84 |
```
|
| 85 |
|
| 86 |
-
> [!
|
| 87 |
-
> Requires a `transformers` version that supports the `DeepseekV4ForCausalLM` architecture. For production
|
| 88 |
|
| 89 |
## Training
|
| 90 |
|
| 91 |
-
Fx-Bio is post-trained on biomedical data (
|
| 92 |
|
| 93 |
`args.json` contains the original Megatron-LM launch arguments for reproducibility.
|
| 94 |
|
|
@@ -99,86 +82,3 @@ For research use only. Model outputs may contain errors or inaccuracies and must
|
|
| 99 |
## License
|
| 100 |
|
| 101 |
[Apache-2.0](LICENSE)
|
| 102 |
-
|
| 103 |
-
---
|
| 104 |
-
|
| 105 |
-
<h1 id="chinese">Fx-Bio (0913) · 中文介绍</h1>
|
| 106 |
-
|
| 107 |
-
**Fx-Bio-0913** 是基于 **DeepSeek-V4-Flash** 进行后训练(post-training)得到的生物领域推理大语言模型,由 The Endless Frontier 实验室发布。模型权重以 **Apache-2.0** 协议开放。
|
| 108 |
-
|
| 109 |
-
## 亮点
|
| 110 |
-
|
| 111 |
-
- 🧬 **领域专精**:在 DeepSeek-V4-Flash 基座之上,使用生物医学推理任务数据(BioMni / BioMystery 相关)进行后训练。
|
| 112 |
-
- 📈 **稳定超越基座**:在两个生物医学推理基准上全面超过 DeepSeek-V4-Flash,困难样本上提升尤为明显——BioMysteryBench Human-Difficult(avg@5)**+5.8**,BiominiBench pass@3 **+6.6**。
|
| 113 |
-
- 📏 **1M 超长上下文**:继承基座模型的 1M token 上下文能力,适配长篇生物医学文献与多跳证据推理。
|
| 114 |
-
- 🔓 **完全开放权重**:bf16 safetensors,Apache-2.0 协议,并附 Megatron-LM 训练启动参数(`args.json`)便于复现。
|
| 115 |
-
|
| 116 |
-
## 评测结果
|
| 117 |
-
|
| 118 |
-
在 **BioMysteryBench** 和 **BiominiBench** 两个生物医学推理基准上与 DeepSeek-V4-Flash 基座模型、Opus-5 的对比如下(Accuracy %,越高越好):
|
| 119 |
-
|
| 120 |
-
<div align="center">
|
| 121 |
-
<img src="assets/fx-bio-eval.png" alt="Fx-Bio 评测结果" width="100%"/>
|
| 122 |
-
</div>
|
| 123 |
-
|
| 124 |
-
**BioMysteryBench**
|
| 125 |
-
|
| 126 |
-
| 指标 | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
|
| 127 |
-
|---|---|---|---|
|
| 128 |
-
| Human-Solvable avg@5 | 85.2 | 89.3 | **85.5** |
|
| 129 |
-
| Human-Difficult avg@5 | 31.8 | 45.9 | **37.6** |
|
| 130 |
-
| pass@5 | 82.2 | 86.7 | **86.7** |
|
| 131 |
-
|
| 132 |
-
**BiominiBench**
|
| 133 |
-
|
| 134 |
-
| 指标 | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
|
| 135 |
-
|---|---|---|---|
|
| 136 |
-
| avg@3 | 71.0 | 79.1 | **77.9** |
|
| 137 |
-
| pass@3 | 78.7 | 86.7 | **85.3** |
|
| 138 |
-
|
| 139 |
-
## 模型信息
|
| 140 |
-
|
| 141 |
-
| 项目 | 内容 |
|
| 142 |
-
|---|---|
|
| 143 |
-
| 基座模型 | DeepSeek-V4-Flash |
|
| 144 |
-
| 模型类型 | MoE 因果语言模型(`DeepseekV4ForCausalLM`) |
|
| 145 |
-
| 权重格式 | bf16 safetensors(114 个分片) |
|
| 146 |
-
| 上下文长度 | 1M tokens |
|
| 147 |
-
| checkpoint | checkpoint-265 |
|
| 148 |
-
| 协议 | Apache-2.0 |
|
| 149 |
-
|
| 150 |
-
## 使用方法
|
| 151 |
-
|
| 152 |
-
```python
|
| 153 |
-
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 154 |
-
|
| 155 |
-
model_id = "endless-frontier/Fx-Bio"
|
| 156 |
-
|
| 157 |
-
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
| 158 |
-
model = AutoModelForCausalLM.from_pretrained(
|
| 159 |
-
model_id,
|
| 160 |
-
torch_dtype="bfloat16",
|
| 161 |
-
device_map="auto",
|
| 162 |
-
)
|
| 163 |
-
|
| 164 |
-
inputs = tokenizer("<你的 prompt>", return_tensors="pt").to(model.device)
|
| 165 |
-
outputs = model.generate(**inputs, max_new_tokens=1024)
|
| 166 |
-
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 167 |
-
```
|
| 168 |
-
|
| 169 |
-
> [!Note]
|
| 170 |
-
> 需要支持 `DeepseekV4ForCausalLM` 架构的 `transformers` 版本。生产环境推理建议使用 vLLM / SGLang 等专用推理框架。
|
| 171 |
-
|
| 172 |
-
## 训练
|
| 173 |
-
|
| 174 |
-
模型在后训练阶段使用了生物医学领域数据(BioMni / BioMystery 相关任务数据)。详细训练配方与数据组成将在后续技术报告中公布。
|
| 175 |
-
|
| 176 |
-
`args.json` 为训练时的 Megatron-LM 启动参数,供复现参考。
|
| 177 |
-
|
| 178 |
-
## 免责声明
|
| 179 |
-
|
| 180 |
-
本模型仅供科研用途。模型输出可能包含错误或不准确信息,不应直接用于临床诊断或医疗决策。
|
| 181 |
-
|
| 182 |
-
## 许可证
|
| 183 |
-
|
| 184 |
-
[Apache-2.0](LICENSE)
|
|
|
|
| 8 |
pipeline_tag: text-generation
|
| 9 |
---
|
| 10 |
|
| 11 |
+
# Fx-Bio
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 12 |
|
| 13 |
+
## Introduction
|
| 14 |
|
| 15 |
+
Fx-Bio-0913 is a biomedical reasoning large language model post-trained on **DeepSeek-V4-Flash**, developed by The Endless Frontier lab. It is specialized for biological and biomedical research tasks — including gene-function puzzles, experimental reasoning, and multi-step evidence integration — through post-training on curated biomedical reasoning data (Biomni / BioMystery-related tasks). The model retains the 1M-token context length of its base model, and its weights are openly released under Apache-2.0.
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
+
## Main Results
|
| 18 |
|
| 19 |
+
Fx-Bio-0913 on two biomedical reasoning benchmarks, **BioMysteryBench** and **BiominiBench**, compared with the DeepSeek-V4-Flash base model and Opus-5. Post-training yields consistent gains over the base model across all metrics, with the largest improvements on harder problems (BioMysteryBench Human-Difficult, +5.8) and on pass@3 of BiominiBench (+6.6).
|
| 20 |
|
| 21 |
+
<p align="center">
|
| 22 |
+
<img src="./assets/fx-bio-eval.png" alt="Fx-Bio-0913 evaluation results on BioMysteryBench and BiominiBench" width="100%">
|
| 23 |
+
</p>
|
|
|
|
|
|
|
| 24 |
|
| 25 |
+
## Benchmark Results
|
|
|
|
|
|
|
|
|
|
|
|
|
| 26 |
|
| 27 |
+
Comparison of Fx-Bio-0913 with representative frontier models on BioMysteryBench (avg@5 and pass@5) and BiominiBench (avg@3 and pass@3). All scores are Accuracy (%) from our internal evaluation pipeline; higher is better.
|
| 28 |
|
| 29 |
+
| Benchmark | Opus<br>5 | GPT<br>5.6 | Gemini 3.8<br>Flash | GLM<br>5.3 | Qwen3.8 Flash<br>Next (1M) | DeepSeek V4.1<br>Flash | DeepSeek V4<br>Flash | Fx-Bio<br>0913 |
|
| 30 |
+
|:--|--:|--:|--:|--:|--:|--:|--:|--:|
|
| 31 |
+
| **BioMysteryBench** | | | | | | | | |
|
| 32 |
+
| Human-Solvable (avg@5) | 89.3 | 85.5 | 88.8 | 84.7 | 86.9 | 89.0 | 85.2 | **85.5** |
|
| 33 |
+
| Human-Difficult (avg@5) | 45.9 | 34.1 | 42.4 | 47.1 | 49.4 | 36.5 | 31.8 | **37.6** |
|
| 34 |
+
| pass@5 | 86.7 | 85.6 | 84.4 | 88.9 | 86.7 | 85.6 | 82.2 | **86.7** |
|
| 35 |
+
| **BiominiBench** | | | | | | | | |
|
| 36 |
+
| avg@3 | 79.1 | 68.2 | 69.3 | 81.1 | 80.6 | 80.9 | 71.0 | **77.9** |
|
| 37 |
+
| pass@3 | 86.7 | 78.6 | 78.5 | 87.2 | 88.9 | 85.9 | 78.7 | **85.3** |
|
| 38 |
|
| 39 |
## Model Overview
|
| 40 |
|
|
|
|
| 66 |
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
|
| 67 |
```
|
| 68 |
|
| 69 |
+
> [!Important]
|
| 70 |
+
> Requires a `transformers` version that supports the `DeepseekV4ForCausalLM` architecture. For production workloads or high-throughput scenarios, a dedicated serving framework (e.g. SGLang or vLLM) with DeepSeek-V4 support is recommended.
|
| 71 |
|
| 72 |
## Training
|
| 73 |
|
| 74 |
+
Fx-Bio is post-trained on biomedical reasoning data (Biomni / BioMystery-related tasks) with Megatron-LM. The detailed training recipe and data composition will be described in an upcoming technical report.
|
| 75 |
|
| 76 |
`args.json` contains the original Megatron-LM launch arguments for reproducibility.
|
| 77 |
|
|
|
|
| 82 |
## License
|
| 83 |
|
| 84 |
[Apache-2.0](LICENSE)
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|