Comment on Why is my home server using so much RAM for cache + buffer?

<- View Parent
Shimitar@downonthestreet.eu ⁨5⁩ ⁨days⁩ ago

Llama CPP can run models offloading with CPU (es MoE models), you have much more control over how you run your models, and overall it’s very much actively developed.

For starters and people without too much willingness to mess up with stuff, ollama is a great choice. Llama.cpp gives you that extra power and flexibility that is so much worth for people who like to tweak and do more.

My personal opinion, of course. But based on having used both and ditched ollama for llama.cpp, so I am also biased, keep in mind.

But I will hardly go back to ollama now :)

original
Sort:hotnewtop