Instructions to use AtomicChat/d1-omni-600M-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use AtomicChat/d1-omni-600M-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf AtomicChat/d1-omni-600M-GGUF:Q4_K_M
Use Docker
docker model run hf.co/AtomicChat/d1-omni-600M-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use AtomicChat/d1-omni-600M-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtomicChat/d1-omni-600M-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtomicChat/d1-omni-600M-GGUF", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AtomicChat/d1-omni-600M-GGUF:Q4_K_M
- Ollama
How to use AtomicChat/d1-omni-600M-GGUF with Ollama:
ollama run hf.co/AtomicChat/d1-omni-600M-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use AtomicChat/d1-omni-600M-GGUF with Docker Model Runner:
docker model run hf.co/AtomicChat/d1-omni-600M-GGUF:Q4_K_M
- Lemonade
How to use AtomicChat/d1-omni-600M-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull AtomicChat/d1-omni-600M-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.d1-omni-600M-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
d1-omni-600M GGUF
Built from Liquid AI's published weights with our own importance matrix, and measured on decisions.
d1-omni-600M is Liquid AI's 587M-parameter decision model. It takes a state (text or JSON, with images or up to 30 seconds of speech) and answers typed questions in one forward pass: yes or no, a pick from named options, or a score. Nothing is generated: the answer is read from the model's scores for the options.
Pick a file
Every number below is measured on one machine. The raw results and logs are in the metrics repo.
same answer: across 1,285 decisions, how often the file gives the same answer as the original weights published in LiquidAI/d1-omni-600M (stored in FP32), run through Liquid's own PyTorch code. Each file answers through llama.cpp's/v1/systemone. The questions are yes/no, choice and score questions over held-out texts in 30 languages and source code. The number of changed answers is in brackets.option drift: the mean total variation distance between the file's option probabilities and the original's. 0 means identical.option KL: the mean KL divergence between the original's option probabilities and the file's.
| File | Size | same answer | option drift | option KL |
|---|---|---|---|---|
BF16 |
764 MB on disk | 99.4% (8) | 0.0029 | 0.00005 |
Q8_0 |
407 MB on disk | 97.7% (29) | 0.0116 | 0.0008 |
AD-Q6_K |
347 MB on disk | 97.4% (33) | 0.0173 | 0.0020 |
AD-Q5_K_M |
293 MB on disk | 95.4% (59) | 0.0324 | 0.0065 |
AD-Q4_K_M |
256 MB on disk | 92.7% (94) | 0.0516 | 0.0153 |
Which one to use. d1-omni is small, and quantization moves its answers more than it moves d1-3B's.
- Q8_0 already changes 2.3% of answers, and 4 bits change 7.3%.
- The largest shift in a single option probability is 0.18 at Q8_0 and 0.46 at 4 bits.
- We recommend
Q8_0, orAD-Q6_K, which is 60 MB smaller and changes only 4 answers more. - Below 6 bits, check the files on your own questions first.
AD- means Atomic Dynamic: the type is chosen per tensor instead of taken from a llama.cpp preset.
- The decision head (the two transformer blocks after the trunk) and the option scorer stay at Q8_0 in every file. They are small, about 26M parameters, and every answer passes through them.
- The token table is a lookup, which an importance matrix does not cover, so it sits above the rest: Q8_0 in
AD-Q6_K, Q6_K inAD-Q5_K_Mand Q5_K inAD-Q4_K_M. - Attention stays at Q8_0, or Q6_K in
AD-Q4_K_M. - The first and last two trunk blocks take one step more than the middle:
ffn_downinAD-Q6_KandAD-Q5_K_M, and the whole feed-forward inAD-Q4_K_M, where the short convolutions also stay at Q5_K. - We compared against llama.cpp's stock
Q4_K_M(247 MB) in the PyTorch measurement described below. The stock file changed 133 answers with an option KL of 0.0255.AD-Q4_K_Mchanged 97, with 44% less KL.
The vision and audio projector comes as mmproj-d1-omni-600M-BF16 and mmproj-d1-omni-600M-Q8_0. We did not
measure image or audio decisions. On one image question, our BF16 with the projector answers like Liquid's
PyTorch code: 0.944 against 0.951 for the same option.
Running it
This needs a llama.cpp build that includes
#30114: commit a657f7e, merged on 8 October, or newer.
d1-omni reads a whole question in one batch, so -b and -ub must hold the longest prompt. 8192 covers an
image with a state of a few thousand tokens.
llama-server -m d1-omni-600M-Q8_0.gguf --mmproj mmproj-d1-omni-600M-Q8_0.gguf -ngl 99 -b 8192 -ub 8192
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "I was charged twice this month, please refund one of them.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
"fraud": "Suspected unauthorised use"}},
"angry": {"type": "noul", "instructions": "Is the customer angry?"}
}
}'
Response from Q8_0, numbers rounded:
{
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.993, "technical": 0.005, "fraud": 0.002}, "confidence": 0.989},
"angry": {"type": "noul", "noul": 0.283}
},
"usage": {"input_tokens": 100, "output_tokens": 0}
}
- Images and audio clips go in a
filesarray as data URLs (data:image/...ordata:audio/...). - The full request format is in the server documentation, and the question schema in the original card.
How these compare to Liquid's own GGUFs
Liquid's GGUFs come in BF16, F16 and Q8_0. Liquid published them on 6 October. On 7 October the repo was offline for a few hours. Later that day Liquid updated the files' metadata for llama.cpp, and each file grew by 32 bytes.
Unlike d1-3B, whose weights Liquid updated on 7 October, d1-omni's weights did not change:
- Before either side updated the metadata, our BF16 file and both of our projector files had the same SHA-256 as Liquid's.
- The metadata matches now too, apart from the model name and tags.
- Liquid did not publish files below 8 bits.
| File | Size | same answer | option drift | option KL |
|---|---|---|---|---|
Liquid Q8_0 |
407 MB | 98.4% (20) | 0.0074 | 0.0003 |
AtomicChat Q8_0 |
407 MB | 98.5% (19) | 0.0076 | 0.0003 |
The two Q8_0 files are the same quantization written by two tools, and they measure the same. These two rows come from the PyTorch measurement described below.
How we made and measured them
- Weights: LiquidAI/d1-omni-600M at revision
414f8d6. When we built the files, llama.cpp had no converter for d1-omni. So each file was written in the layout of Liquid's GGUF: the same tensor names, order and metadata, including the decision type, the temperatures and thesystemonetemplate.- Every tensor was matched against the checkpoint before writing.
- The audio encoder's batch norms are folded into a scale and a shift, as llama.cpp's converters do.
- The converter that came with #30114 writes the same tensors.
- The text model is identical.
- In the projector, the folded batch norms differ only in the last bit of a float32.
- Quantization: llama.cpp's
llama-quantizeat commit18b5f8b. The projector's Q8_0 is written the way Liquid wrote theirs. - Importance matrix: computed in PyTorch, since llama.cpp cannot run the model, and saved in llama.cpp's
format.
- Inputs were collected at the 92 linear layers of the trunk.
- The prompts were 2,378 decision prompts, about 1.5M tokens. Their states come from the calibration corpora pool: Wikipedia in 30 languages, code and structured files. Each prompt carries yes/no, choice and score questions.
- Metadata: on 8 October, after #30114 was merged, the text files were re-stamped with
lfm2.decision.type = lfm2-d1-omniand the converter'ssystemonetemplate. No tensor changed. The projectors needed no change. - Decisions: 1,285 questions over the held-out calib-corpora
eval/neutralandeval/codetexts, none of them in the calibration set. The reference is Liquid's own PyTorch code in FP32.- The table above comes from llama.cpp
a657f7eon/v1/systemone, on an NVIDIA A10. - Before the runtime existed, each file was also read back, dequantized, into Liquid's PyTorch model. That
measures only the error the quantized weights add. It gave 19, 30, 63 and 97 changed answers from
Q8_0down toAD-Q4_K_M, which agrees with the table. - Both sets of results are in the metrics repo.
- The table above comes from llama.cpp
Model details
- Base: LiquidAI/d1-omni-600M, a decision model post-trained from LFM2.5-Encoder-350M.
- Size: 587M parameters in all: a 381M trunk and decision head, a 94M SigLIP2 vision encoder and a 112M FastConformer audio encoder.
- Context and vocabulary: 16,384 positions for text, image and audio together, and a 65,536-token vocabulary.
- License: LFM Open License v1.0; see
LICENSE.
- Downloads last month
- 328
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for AtomicChat/d1-omni-600M-GGUF
Base model
LiquidAI/LFM2.5-350M-Base

