Instructions to use JetBrains/Mellum2-12B-A2.5B-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JetBrains/Mellum2-12B-A2.5B-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JetBrains/Mellum2-12B-A2.5B-Base")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JetBrains/Mellum2-12B-A2.5B-Base") model = AutoModelForCausalLM.from_pretrained("JetBrains/Mellum2-12B-A2.5B-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JetBrains/Mellum2-12B-A2.5B-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JetBrains/Mellum2-12B-A2.5B-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2-12B-A2.5B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/JetBrains/Mellum2-12B-A2.5B-Base
- SGLang
How to use JetBrains/Mellum2-12B-A2.5B-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JetBrains/Mellum2-12B-A2.5B-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2-12B-A2.5B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JetBrains/Mellum2-12B-A2.5B-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JetBrains/Mellum2-12B-A2.5B-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use JetBrains/Mellum2-12B-A2.5B-Base with Docker Model Runner:
docker model run hf.co/JetBrains/Mellum2-12B-A2.5B-Base
Add logo, use case note, vLLM serving, and Quickstart
Browse files- README.md +36 -0
- mellum-logo-dark.svg +40 -0
- mellum-logo.svg +40 -0
README.md
CHANGED
|
@@ -206,8 +206,16 @@ model-index:
|
|
| 206 |
license: apache-2.0
|
| 207 |
---
|
| 208 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 209 |
# Mellum 2 Base
|
| 210 |
|
|
|
|
|
|
|
|
|
|
| 211 |
## Mellum 2 Base Highlights
|
| 212 |
|
| 213 |
Mellum 2 Base is a long-context pretrained causal language model trained by JetBrains.
|
|
@@ -245,6 +253,34 @@ This repository contains one checkpoint from the Mellum 2 family.
|
|
| 245 |
- Vocabulary Size: 98,304
|
| 246 |
- Precision: bfloat16
|
| 247 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 248 |
## Evaluation
|
| 249 |
|
| 250 |
Evaluation results are available in the model card. All values are self-reported by JetBrains.
|
|
|
|
| 206 |
license: apache-2.0
|
| 207 |
---
|
| 208 |
|
| 209 |
+
<picture>
|
| 210 |
+
<source media="(prefers-color-scheme: dark)" srcset="mellum-logo-dark.svg">
|
| 211 |
+
<img alt="Mellum" src="mellum-logo.svg" width="320">
|
| 212 |
+
</picture>
|
| 213 |
+
|
| 214 |
# Mellum 2 Base
|
| 215 |
|
| 216 |
+
> [!Note]
|
| 217 |
+
> Use this checkpoint as the starting point for your own fine-tuning, alignment, or domain adaptation on top of the long-context base. For instruction-following or reasoning tasks out of the box, use [Instruct](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct) or [Thinking](https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking) instead.
|
| 218 |
+
|
| 219 |
## Mellum 2 Base Highlights
|
| 220 |
|
| 221 |
Mellum 2 Base is a long-context pretrained causal language model trained by JetBrains.
|
|
|
|
| 253 |
- Vocabulary Size: 98,304
|
| 254 |
- Precision: bfloat16
|
| 255 |
|
| 256 |
+
## Serving with vLLM
|
| 257 |
+
|
| 258 |
+
```sh
|
| 259 |
+
vllm serve JetBrains/Mellum2-12B-A2.5B-Base --max-model-len 131072
|
| 260 |
+
```
|
| 261 |
+
|
| 262 |
+
## Quickstart
|
| 263 |
+
|
| 264 |
+
Text-Only Input (base model — use the completions endpoint, not chat)
|
| 265 |
+
|
| 266 |
+
```python
|
| 267 |
+
from openai import OpenAI
|
| 268 |
+
# Configured by environment variables
|
| 269 |
+
client = OpenAI()
|
| 270 |
+
|
| 271 |
+
completion = client.completions.create(
|
| 272 |
+
model="JetBrains/Mellum2-12B-A2.5B-Base",
|
| 273 |
+
prompt="def fibonacci(n):\n ",
|
| 274 |
+
max_tokens=81920,
|
| 275 |
+
temperature=0.6,
|
| 276 |
+
top_p=0.95,
|
| 277 |
+
extra_body={
|
| 278 |
+
"top_k": 20,
|
| 279 |
+
},
|
| 280 |
+
)
|
| 281 |
+
print("Completion:", completion)
|
| 282 |
+
```
|
| 283 |
+
|
| 284 |
## Evaluation
|
| 285 |
|
| 286 |
Evaluation results are available in the model card. All values are self-reported by JetBrains.
|
mellum-logo-dark.svg
ADDED
|
|
mellum-logo.svg
ADDED
|
|