Please don't forget about us 35B MOE users!

#70
by CYISNOTHERE - opened

Coping

35B 35B 35B!)

I am absolutely loving the 27b model. 45% faster token generation. From 26tokens a second on 3.6 27b to 41 tokens a second on 3.8 27b. And thinking times are much faster from 191 seconds to 119 seconds. With a 5950x 64gb ddr 4 3600mhz ram and dual 5060ti 16gb each. With ollama front end. So far im happy with it.

I am absolutely loving the 27b model. 45% faster token generation. From 26tokens a second on 3.6 27b to 41 tokens a second on 3.8 27b. And thinking times are much faster from 191 seconds to 119 seconds. With a 5950x 64gb ddr 4 3600mhz ram and dual 5060ti 16gb each. With ollama front end. So far im happy with it.

What's the config/quant? The KV offload to RAM should make those numbers a lot lower on your hardware for a 27B dense model.

Sign up or log in to comment