PetteriPano@lemmy.world 4 hours ago
I try most models, that get support in llama.cpp or forks. IIRC I wasn’t overly impressed with this one. Slower than it should be, mostly because its KV cache is huge.
PetteriPano@lemmy.world 4 hours ago
I try most models, that get support in llama.cpp or forks. IIRC I wasn’t overly impressed with this one. Slower than it should be, mostly because its KV cache is huge.
surewhynotlem@lemmy.world 1 hour ago
I keep falling back to qwen. Though bonsai did give it a good run for a while.
But I only have an old 8gb card.
greybeard@feddit.online 38 minutes ago
Qwen is what I’ve found to work the best for me so far. I haven’t done an exhaustive search or anything, but the mixture of experts 35b models seem to work pretty well.
I’ve got a little more vRAM to play with, 20GB. It still struggles with giving it enough context to be useful for agentic stuff.