Comment on Selfhosting Sunday! What's up?
brucethemoose@lemmy.world 1 day ago
I know this isn’t everyone’s cup of tea, but I’m excited about Deepseek V4 flash. It’s (for me) the perfect size and architecture to self-host and LLM.
My box is chugging through a queue:
-
Figure out why my swap is going crazy, and how to ban processes from it [Done].
-
Make an ik_llama.cpp iMatrix for Deepseek V4 [Done].
-
Figure out what quantization isn’t working [Done].
-
Make a test IQ2_KL/MXFP4_R8 quant to see how it does squeezed onto my box [in progress].
-
Test. Inevitably troubleshoot the dozen other things that go wrong. Figure out how much spare RAM that leaves me.
-
Make a higher quality IQ3_KT quantization. This will take all night.
-
KLD test both of them vs the full precision, to quantify quantization loss. Likely an overnight task, too.
-
Try merging the new model release with the base model, 50/50, for a less “deep fried” model. imatrix, quant, test.
What are the advantages of Deepseek V4 flash?
brucethemoose@lemmy.world 1 day ago
It’s just under 300B, trained at FP4; absolutely the perfect size for servers with 128GB-192GB CPU RAM to spare.
Its fast. I’m getting 17 tokens/sec on a single RTX 3090 GPU, all experts offloaded to RAM; for a 300B model this smart, that’s crazy fast.
Its attention mechanism is cutting edge, good for long context without too much processing time.
The benchmarks for coding/agenic usage are absolutely bonkers, within margin of error of frontier models or Deepseek Pro in some cases: huggingface.co/…/DeepSeek-V4-Flash-0731
…Though I don’t put much stock in benchmarks. And I haven’t tested it enough to tell you if it lives up to that hype for specific use cases.
Deepseek also publishes its base model. That means I can “unfry” the model with a merge if I have to.
I am afraid of the the model being “overfit” to coding and agenic stuff.
For reference, my previous favorite model was Xiaomi MiMo V2.5 310B. It benched well, but it also smelt “smart” outside of benchmarks, like in knowledge of trivia without tool usage/internet access or comprehension of weird questions.
Sweet.