Comment on how do you manage your server?
something183786@lemmy.world 2 weeks agoWhat models are you using? What hardware are you running them on? I’m curious if I can replicate your success
Comment on how do you manage your server?
something183786@lemmy.world 2 weeks agoWhat models are you using? What hardware are you running them on? I’m curious if I can replicate your success
9tr6gyp3@lemmy.world 2 weeks ago
Hardware I’m running:
I’m mainly using these MoE models:
Qwen-3.6-35B-A3B
Q4_K_M quant quality
128k context (conversation length before it compacts)
Gives me about 260 prefill and 19 token gen speeds
Gemma4-26B-A4B
Q8 quant quality
128k context length
Gives me about 190 prefill and 14 token gen speeds