Model card?
rec settings? what was your goal for this one?
It most certainly is Gemma-4. I got it running comfortably with the following:
vllm serve ~/ai/models/Orion-26B-A4B-v1
--quantization fp8
--tool-call-parser gemma4
--enable-auto-tool-choice
--max-model-len 262144
--gpu-memory-utilization 0.90
--served-model-name Orion-26B-A4B-v1
--max-num-seqs 2
--kv-cache-dtype fp8
--reasoning-parser gemma4
--chat-template ~/ai/models/Orion-26B-A4B-v1/chat_template.jinja
Pros: lightning fast, proper tools usage (Ive noticed almost no degradation - very good). uncensored - no refusals yet.
Cons: As always with 3-4AB you need to fight with the model for it to do what you actually want. Prose is not outstanding. Not fit for deep analysis - not that attentive to details. Indecisive - one of the hardest models to force doing anything automatically. It is pretty prone to get 10 "I am ready to proceed with the next step" confirmations even with AGENTS.md having fully fheshed out proccess. I am still suspecting AGENTS.md itself as the main culprit, but I never had the same problem with other models.
I've yet to try finding thinking mode, but I suspect its without one.
Note: my use scenarious are pretty specific, so I am probably completely ignoring this models intended strengths =D
I've yet to try finding thinking mode, but I suspect its without one.
It definitely can reason, witnessed it myself