Comment

Comment on Self-Hosted AI is pretty darn cool

Have you found much practical use for small models yet? I love the idea that even the 1.1B tinyllama model can run on my phone, but haven’t found much real world use for it yet. Llama3 8b feels better, but not much better for even emails as it’s a bit dumb

source

Sort:hotnew top

chagall@lemmy.world ⁨5⁩ ⁨months⁩ ago
I use my phone all the time, but I just use a wireguard VPN to tunnel into my home container of Open WebUI. Then I can interact with my desktop machine using a NVIDIA gpu. I’m currently testing mistral-nemo. It’s pretty great but it gets a bit verbose sometimes.

source
- kureta@lemmy.ml ⁨5⁩ ⁨months⁩ ago
  I am also using open webui. Most LLMs are too verbose for me, so I created a model in open-webui with system prompt “Do not repeat the questions. Avoid giving lists as answers. Do not summarize the answer at the end. If asked a follow-up question, respond with only new information, do not repeat previously stated information.” and named it No Nonsense.
  
  source
  - kate@lemmy.uhhoh.com ⁨5⁩ ⁨months⁩ ago
    for some reason chatgpt responds well to “no yapping”
    
    source
  - chagall@lemmy.world ⁨5⁩ ⁨months⁩ ago
    That’s really smart. I just found out about fabric yesterday and it is helping me with things like what you stated. Prompt engineering is a huge thing.
    
    source
coffee_with_cream@sh.itjust.works ⁨5⁩ ⁨months⁩ ago
Imo it’s worthwhile to just run the biggest model available and rent expensive GPU time. It still amounts to very little overall and you get much better results. Project dependent of course

source