Comment on how do you manage your server?
9tr6gyp3@lemmy.world 2 weeks agoHardware I’m running:
- 8/16 AMD CPU
- 32GB system RAM
- 12GB GPU VRAM
I’m mainly using these MoE models:
-
Qwen-3.6-35B-A3B
-
Q4_K_M quant quality
-
128k context (conversation length before it compacts)
-
Gives me about 260 prefill and 19 token gen speeds
-
Gemma4-26B-A4B
-
Q8 quant quality
-
128k context length
-
Gives me about 190 prefill and 14 token gen speeds