Qwen is what I’ve found to work the best for me so far. I haven’t done an exhaustive search or anything, but the mixture of experts 35b models seem to work pretty well.
I’ve got a little more vRAM to play with, 20GB. It still struggles with giving it enough context to be useful for agentic stuff.
I have a bit more memory, so I jump between the two qwen 3.8s.
Glimmer muse was a good contender in the MoE space for a few days. Might be more viable on 8gb.