d1-omni-600M GGUF

Built from Liquid AI's published weights with our own importance matrix, and measured on decisions.

Atomic Chat Discord GitHub

d1-omni-600M is Liquid AI's 587M-parameter decision model. It takes a state (text or JSON, with images or up to 30 seconds of speech) and answers typed questions in one forward pass: yes or no, a pick from named options, or a score. Nothing is generated: the answer is read from the model's scores for the options.

Pick a file

Every number below is measured on one machine. The raw results and logs are in the metrics repo.

  • same answer: across 1,285 decisions, how often the file gives the same answer as the original weights published in LiquidAI/d1-omni-600M (stored in FP32), run through Liquid's own PyTorch code. Each file answers through llama.cpp's /v1/systemone. The questions are yes/no, choice and score questions over held-out texts in 30 languages and source code. The number of changed answers is in brackets.
  • option drift: the mean total variation distance between the file's option probabilities and the original's. 0 means identical.
  • option KL: the mean KL divergence between the original's option probabilities and the file's.
File Size same answer option drift option KL
BF16 764 MB on disk 99.4% (8) 0.0029 0.00005
Q8_0 407 MB on disk 97.7% (29) 0.0116 0.0008
AD-Q6_K 347 MB on disk 97.4% (33) 0.0173 0.0020
AD-Q5_K_M 293 MB on disk 95.4% (59) 0.0324 0.0065
AD-Q4_K_M 256 MB on disk 92.7% (94) 0.0516 0.0153

Which one to use. d1-omni is small, and quantization moves its answers more than it moves d1-3B's.

  • Q8_0 already changes 2.3% of answers, and 4 bits change 7.3%.
  • The largest shift in a single option probability is 0.18 at Q8_0 and 0.46 at 4 bits.
  • We recommend Q8_0, or AD-Q6_K, which is 60 MB smaller and changes only 4 answers more.
  • Below 6 bits, check the files on your own questions first.

AD- means Atomic Dynamic: the type is chosen per tensor instead of taken from a llama.cpp preset.

  • The decision head (the two transformer blocks after the trunk) and the option scorer stay at Q8_0 in every file. They are small, about 26M parameters, and every answer passes through them.
  • The token table is a lookup, which an importance matrix does not cover, so it sits above the rest: Q8_0 in AD-Q6_K, Q6_K in AD-Q5_K_M and Q5_K in AD-Q4_K_M.
  • Attention stays at Q8_0, or Q6_K in AD-Q4_K_M.
  • The first and last two trunk blocks take one step more than the middle: ffn_down in AD-Q6_K and AD-Q5_K_M, and the whole feed-forward in AD-Q4_K_M, where the short convolutions also stay at Q5_K.
  • We compared against llama.cpp's stock Q4_K_M (247 MB) in the PyTorch measurement described below. The stock file changed 133 answers with an option KL of 0.0255. AD-Q4_K_M changed 97, with 44% less KL.

The vision and audio projector comes as mmproj-d1-omni-600M-BF16 and mmproj-d1-omni-600M-Q8_0. We did not measure image or audio decisions. On one image question, our BF16 with the projector answers like Liquid's PyTorch code: 0.944 against 0.951 for the same option.

Running it

This needs a llama.cpp build that includes #30114: commit a657f7e, merged on 8 October, or newer. d1-omni reads a whole question in one batch, so -b and -ub must hold the longest prompt. 8192 covers an image with a state of a few thousand tokens.

llama-server -m d1-omni-600M-Q8_0.gguf --mmproj mmproj-d1-omni-600M-Q8_0.gguf -ngl 99 -b 8192 -ub 8192
curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "I was charged twice this month, please refund one of them.",
  "questions": {
    "team": {"type": "choice", "instructions": "Which team should handle this?",
             "criteria": {"billing": "Charges, refunds, invoices", "technical": "App or site faults",
                          "fraud": "Suspected unauthorised use"}},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"}
  }
}'

Response from Q8_0, numbers rounded:

{
  "answers": {
    "team": {"type": "choice", "choice": "billing",
             "probabilities": {"billing": 0.993, "technical": 0.005, "fraud": 0.002}, "confidence": 0.989},
    "angry": {"type": "noul", "noul": 0.283}
  },
  "usage": {"input_tokens": 100, "output_tokens": 0}
}
  • Images and audio clips go in a files array as data URLs (data:image/... or data:audio/...).
  • The full request format is in the server documentation, and the question schema in the original card.

How these compare to Liquid's own GGUFs

Liquid's GGUFs come in BF16, F16 and Q8_0. Liquid published them on 6 October. On 7 October the repo was offline for a few hours. Later that day Liquid updated the files' metadata for llama.cpp, and each file grew by 32 bytes.

Unlike d1-3B, whose weights Liquid updated on 7 October, d1-omni's weights did not change:

  • Before either side updated the metadata, our BF16 file and both of our projector files had the same SHA-256 as Liquid's.
  • The metadata matches now too, apart from the model name and tags.
  • Liquid did not publish files below 8 bits.
File Size same answer option drift option KL
Liquid Q8_0 407 MB 98.4% (20) 0.0074 0.0003
AtomicChat Q8_0 407 MB 98.5% (19) 0.0076 0.0003

The two Q8_0 files are the same quantization written by two tools, and they measure the same. These two rows come from the PyTorch measurement described below.

How we made and measured them

  • Weights: LiquidAI/d1-omni-600M at revision 414f8d6. When we built the files, llama.cpp had no converter for d1-omni. So each file was written in the layout of Liquid's GGUF: the same tensor names, order and metadata, including the decision type, the temperatures and the systemone template.
    • Every tensor was matched against the checkpoint before writing.
    • The audio encoder's batch norms are folded into a scale and a shift, as llama.cpp's converters do.
    • The converter that came with #30114 writes the same tensors.
      • The text model is identical.
      • In the projector, the folded batch norms differ only in the last bit of a float32.
  • Quantization: llama.cpp's llama-quantize at commit 18b5f8b. The projector's Q8_0 is written the way Liquid wrote theirs.
  • Importance matrix: computed in PyTorch, since llama.cpp cannot run the model, and saved in llama.cpp's format.
    • Inputs were collected at the 92 linear layers of the trunk.
    • The prompts were 2,378 decision prompts, about 1.5M tokens. Their states come from the calibration corpora pool: Wikipedia in 30 languages, code and structured files. Each prompt carries yes/no, choice and score questions.
  • Metadata: on 8 October, after #30114 was merged, the text files were re-stamped with lfm2.decision.type = lfm2-d1-omni and the converter's systemone template. No tensor changed. The projectors needed no change.
  • Decisions: 1,285 questions over the held-out calib-corpora eval/neutral and eval/code texts, none of them in the calibration set. The reference is Liquid's own PyTorch code in FP32.
    • The table above comes from llama.cpp a657f7e on /v1/systemone, on an NVIDIA A10.
    • Before the runtime existed, each file was also read back, dequantized, into Liquid's PyTorch model. That measures only the error the quantized weights add. It gave 19, 30, 63 and 97 changed answers from Q8_0 down to AD-Q4_K_M, which agrees with the table.
    • Both sets of results are in the metrics repo.

Model details

  • Base: LiquidAI/d1-omni-600M, a decision model post-trained from LFM2.5-Encoder-350M.
  • Size: 587M parameters in all: a 381M trunk and decision head, a 94M SigLIP2 vision encoder and a 112M FastConformer audio encoder.
  • Context and vocabulary: 16,384 positions for text, image and audio together, and a 65,536-token vocabulary.
  • License: LFM Open License v1.0; see LICENSE.
Downloads last month
328
GGUF
Model size
0.4B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AtomicChat/d1-omni-600M-GGUF

Quantized
(10)
this model

Collection including AtomicChat/d1-omni-600M-GGUF