Model card?

#2
by X-Wave - opened

rec settings? what was your goal for this one?

It most certainly is Gemma-4. I got it running comfortably with the following:

vllm serve ~/ai/models/Orion-26B-A4B-v1
--quantization fp8
--tool-call-parser gemma4
--enable-auto-tool-choice
--max-model-len 262144
--gpu-memory-utilization 0.90
--served-model-name Orion-26B-A4B-v1
--max-num-seqs 2
--kv-cache-dtype fp8
--reasoning-parser gemma4
--chat-template ~/ai/models/Orion-26B-A4B-v1/chat_template.jinja

Pros: lightning fast, proper tools usage (Ive noticed almost no degradation - very good). uncensored - no refusals yet.
Cons: As always with 3-4AB you need to fight with the model for it to do what you actually want. Prose is not outstanding. Not fit for deep analysis - not that attentive to details. Indecisive - one of the hardest models to force doing anything automatically. It is pretty prone to get 10 "I am ready to proceed with the next step" confirmations even with AGENTS.md having fully fheshed out proccess. I am still suspecting AGENTS.md itself as the main culprit, but I never had the same problem with other models.

I've yet to try finding thinking mode, but I suspect its without one.

Note: my use scenarious are pretty specific, so I am probably completely ignoring this models intended strengths =D

I've yet to try finding thinking mode, but I suspect its without one.

It definitely can reason, witnessed it myself

X-Wave changed discussion status to closed
X-Wave changed discussion status to open
X-Wave changed discussion status to closed
TheDrummer changed discussion status to open

Sign up or log in to comment