muskliu commited on
Commit
031e9c8
·
verified ·
1 Parent(s): aa88f43

Rewrite model card: BigBang-style structure, expanded benchmark comparison, remove banner

Browse files
Files changed (1) hide show
  1. README.md +22 -122
README.md CHANGED
@@ -8,50 +8,33 @@ tags:
8
  pipeline_tag: text-generation
9
  ---
10
 
11
- <div align="center">
12
- <img src="assets/fx-bio-banner.png" alt="Fx-Bio Banner" width="100%"/>
13
-
14
- <p>
15
- <a href="https://huggingface.co/endless-frontier/Fx-Bio"><img alt="Model" src="https://img.shields.io/badge/%F0%9F%A4%97%20Model-endless--frontier%2FFx--Bio-536af5"/></a>
16
- <a href="LICENSE"><img alt="License" src="https://img.shields.io/badge/License-Apache--2.0-blue.svg"/></a>
17
- <img alt="Base Model" src="https://img.shields.io/badge/Base-DeepSeek--V4--Flash-6b5ce7"/>
18
- <img alt="Context" src="https://img.shields.io/badge/Context-1M%20tokens-0d9488"/>
19
- </p>
20
-
21
- <p>English ⬇️ &nbsp;|&nbsp; <a href="#chinese">中文 ⬇️</a></p>
22
- </div>
23
-
24
- **Fx-Bio-0913** is a biomedical reasoning large language model post-trained on **DeepSeek-V4-Flash**, developed by The Endless Frontier lab. Model weights are openly released under **Apache-2.0**.
25
 
26
- ## Highlights
27
 
28
- - 🧬 **Domain-specialized**: post-trained on biomedical reasoning tasks (BioMni / BioMystery-related data) on top of DeepSeek-V4-Flash.
29
- - 📈 **Consistent gains over the base model** on biomedical reasoning benchmarks, especially on harder problems — e.g. **+5.8** on BioMysteryBench Human-Difficult (avg@5) and **+6.6** on BiominiBench pass@3 over DeepSeek-V4-Flash.
30
- - 📏 **1M-token context length**, inherited from the base model, for long biomedical documents and multi-hop evidence.
31
- - 🔓 **Fully open weights** (bf16 safetensors) under Apache-2.0, with Megatron-LM training arguments (`args.json`) included for reproducibility.
32
 
33
- ## Performance
34
 
35
- Evaluated on **BioMysteryBench** and **BiominiBench** against the DeepSeek-V4-Flash base model and Opus-5 (Accuracy %, higher is better):
36
 
37
- <div align="center">
38
- <img src="assets/fx-bio-eval.png" alt="Fx-Bio Evaluation Results" width="100%"/>
39
- </div>
40
-
41
- **BioMysteryBench**
42
 
43
- | Metric | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
44
- |---|---|---|---|
45
- | Human-Solvable avg@5 | 85.2 | 89.3 | **85.5** |
46
- | Human-Difficult avg@5 | 31.8 | 45.9 | **37.6** |
47
- | pass@5 | 82.2 | 86.7 | **86.7** |
48
 
49
- **BiominiBench**
50
 
51
- | Metric | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
52
- |---|---|---|---|
53
- | avg@3 | 71.0 | 79.1 | **77.9** |
54
- | pass@3 | 78.7 | 86.7 | **85.3** |
 
 
 
 
 
55
 
56
  ## Model Overview
57
 
@@ -83,12 +66,12 @@ outputs = model.generate(**inputs, max_new_tokens=1024)
83
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
84
  ```
85
 
86
- > [!Note]
87
- > Requires a `transformers` version that supports the `DeepseekV4ForCausalLM` architecture. For production inference, a dedicated serving framework (e.g. vLLM / SGLang) is recommended.
88
 
89
  ## Training
90
 
91
- Fx-Bio is post-trained on biomedical data (BioMni / BioMystery-related tasks). The detailed training recipe and data composition will be described in an upcoming technical report.
92
 
93
  `args.json` contains the original Megatron-LM launch arguments for reproducibility.
94
 
@@ -99,86 +82,3 @@ For research use only. Model outputs may contain errors or inaccuracies and must
99
  ## License
100
 
101
  [Apache-2.0](LICENSE)
102
-
103
- ---
104
-
105
- <h1 id="chinese">Fx-Bio (0913) · 中文介绍</h1>
106
-
107
- **Fx-Bio-0913** 是基于 **DeepSeek-V4-Flash** 进行后训练(post-training)得到的生物领域推理大语言模型,由 The Endless Frontier 实验室发布。模型权重以 **Apache-2.0** 协议开放。
108
-
109
- ## 亮点
110
-
111
- - 🧬 **领域专精**:在 DeepSeek-V4-Flash 基座之上,使用生物医学推理任务数据(BioMni / BioMystery 相关)进行后训练。
112
- - 📈 **稳定超越基座**:在两个生物医学推理基准上全面超过 DeepSeek-V4-Flash,困难样本上提升尤为明显——BioMysteryBench Human-Difficult(avg@5)**+5.8**,BiominiBench pass@3 **+6.6**。
113
- - 📏 **1M 超长上下文**:继承基座模型的 1M token 上下文能力,适配长篇生物医学文献与多跳证据推理。
114
- - 🔓 **完全开放权重**:bf16 safetensors,Apache-2.0 协议,并附 Megatron-LM 训练启动参数(`args.json`)便于复现。
115
-
116
- ## 评测结果
117
-
118
- 在 **BioMysteryBench** 和 **BiominiBench** 两个生物医学推理基准上与 DeepSeek-V4-Flash 基座模型、Opus-5 的对比如下(Accuracy %,越高越好):
119
-
120
- <div align="center">
121
- <img src="assets/fx-bio-eval.png" alt="Fx-Bio 评测结果" width="100%"/>
122
- </div>
123
-
124
- **BioMysteryBench**
125
-
126
- | 指标 | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
127
- |---|---|---|---|
128
- | Human-Solvable avg@5 | 85.2 | 89.3 | **85.5** |
129
- | Human-Difficult avg@5 | 31.8 | 45.9 | **37.6** |
130
- | pass@5 | 82.2 | 86.7 | **86.7** |
131
-
132
- **BiominiBench**
133
-
134
- | 指标 | DeepSeek-V4-Flash | Opus-5 | **Fx-Bio-0913** |
135
- |---|---|---|---|
136
- | avg@3 | 71.0 | 79.1 | **77.9** |
137
- | pass@3 | 78.7 | 86.7 | **85.3** |
138
-
139
- ## 模型信息
140
-
141
- | 项目 | 内容 |
142
- |---|---|
143
- | 基座模型 | DeepSeek-V4-Flash |
144
- | 模型类型 | MoE 因果语言模型(`DeepseekV4ForCausalLM`) |
145
- | 权重格式 | bf16 safetensors(114 个分片) |
146
- | 上下文长度 | 1M tokens |
147
- | checkpoint | checkpoint-265 |
148
- | 协议 | Apache-2.0 |
149
-
150
- ## 使用方法
151
-
152
- ```python
153
- from transformers import AutoModelForCausalLM, AutoTokenizer
154
-
155
- model_id = "endless-frontier/Fx-Bio"
156
-
157
- tokenizer = AutoTokenizer.from_pretrained(model_id)
158
- model = AutoModelForCausalLM.from_pretrained(
159
- model_id,
160
- torch_dtype="bfloat16",
161
- device_map="auto",
162
- )
163
-
164
- inputs = tokenizer("<你的 prompt>", return_tensors="pt").to(model.device)
165
- outputs = model.generate(**inputs, max_new_tokens=1024)
166
- print(tokenizer.decode(outputs[0], skip_special_tokens=True))
167
- ```
168
-
169
- > [!Note]
170
- > 需要支持 `DeepseekV4ForCausalLM` 架构的 `transformers` 版本。生产环境推理建议使用 vLLM / SGLang 等专用推理框架。
171
-
172
- ## 训练
173
-
174
- 模型在后训练阶段使用了生物医学领域数据(BioMni / BioMystery 相关任务数据)。详细训练配方与数据组成将在后续技术报告中公布。
175
-
176
- `args.json` 为训练时的 Megatron-LM 启动参数,供复现参考。
177
-
178
- ## 免责声明
179
-
180
- 本模型仅供科研用途。模型输出可能包含错误或不准确信息,不应直接用于临床诊断或医疗决策。
181
-
182
- ## 许可证
183
-
184
- [Apache-2.0](LICENSE)
 
8
  pipeline_tag: text-generation
9
  ---
10
 
11
+ # Fx-Bio
 
 
 
 
 
 
 
 
 
 
 
 
 
12
 
13
+ ## Introduction
14
 
15
+ Fx-Bio-0913 is a biomedical reasoning large language model post-trained on **DeepSeek-V4-Flash**, developed by The Endless Frontier lab. It is specialized for biological and biomedical research tasks — including gene-function puzzles, experimental reasoning, and multi-step evidence integration — through post-training on curated biomedical reasoning data (Biomni / BioMystery-related tasks). The model retains the 1M-token context length of its base model, and its weights are openly released under Apache-2.0.
 
 
 
16
 
17
+ ## Main Results
18
 
19
+ Fx-Bio-0913 on two biomedical reasoning benchmarks, **BioMysteryBench** and **BiominiBench**, compared with the DeepSeek-V4-Flash base model and Opus-5. Post-training yields consistent gains over the base model across all metrics, with the largest improvements on harder problems (BioMysteryBench Human-Difficult, +5.8) and on pass@3 of BiominiBench (+6.6).
20
 
21
+ <p align="center">
22
+ <img src="./assets/fx-bio-eval.png" alt="Fx-Bio-0913 evaluation results on BioMysteryBench and BiominiBench" width="100%">
23
+ </p>
 
 
24
 
25
+ ## Benchmark Results
 
 
 
 
26
 
27
+ Comparison of Fx-Bio-0913 with representative frontier models on BioMysteryBench (avg@5 and pass@5) and BiominiBench (avg@3 and pass@3). All scores are Accuracy (%) from our internal evaluation pipeline; higher is better.
28
 
29
+ | Benchmark | Opus<br>5 | GPT<br>5.6 | Gemini 3.8<br>Flash | GLM<br>5.3 | Qwen3.8 Flash<br>Next (1M) | DeepSeek V4.1<br>Flash | DeepSeek V4<br>Flash | Fx-Bio<br>0913 |
30
+ |:--|--:|--:|--:|--:|--:|--:|--:|--:|
31
+ | **BioMysteryBench** | | | | | | | | |
32
+ | Human-Solvable (avg@5) | 89.3 | 85.5 | 88.8 | 84.7 | 86.9 | 89.0 | 85.2 | **85.5** |
33
+ | Human-Difficult (avg@5) | 45.9 | 34.1 | 42.4 | 47.1 | 49.4 | 36.5 | 31.8 | **37.6** |
34
+ | pass@5 | 86.7 | 85.6 | 84.4 | 88.9 | 86.7 | 85.6 | 82.2 | **86.7** |
35
+ | **BiominiBench** | | | | | | | | |
36
+ | avg@3 | 79.1 | 68.2 | 69.3 | 81.1 | 80.6 | 80.9 | 71.0 | **77.9** |
37
+ | pass@3 | 86.7 | 78.6 | 78.5 | 87.2 | 88.9 | 85.9 | 78.7 | **85.3** |
38
 
39
  ## Model Overview
40
 
 
66
  print(tokenizer.decode(outputs[0], skip_special_tokens=True))
67
  ```
68
 
69
+ > [!Important]
70
+ > Requires a `transformers` version that supports the `DeepseekV4ForCausalLM` architecture. For production workloads or high-throughput scenarios, a dedicated serving framework (e.g. SGLang or vLLM) with DeepSeek-V4 support is recommended.
71
 
72
  ## Training
73
 
74
+ Fx-Bio is post-trained on biomedical reasoning data (Biomni / BioMystery-related tasks) with Megatron-LM. The detailed training recipe and data composition will be described in an upcoming technical report.
75
 
76
  `args.json` contains the original Megatron-LM launch arguments for reproducibility.
77
 
 
82
  ## License
83
 
84
  [Apache-2.0](LICENSE)