Comment

Comment on Grok praises Hitler, gives credit to Musk for removing “woke filters”

brucethemoose@lemmy.world ⁨10⁩ ⁨months⁩ ago

Nitpick: it was never ‘filtered’

LLMs can be trained to refuse excessively (which is kinda stupid and is objectively proven to make them dumber), but the correct term is ‘biased’. If it was filtered, it would literally give empty responses for anything deemed harmful, or at least noticably take some time to retry.

They trained it to praise hitler, intentionally. They didn’t remove any guardrails.

source

Sort:hotnew top

TheFogan@programming.dev ⁨10⁩ ⁨months⁩ ago

They trained it to praise hitler, intentionally. They didn’t remove any guardrails. Not that Musk acolytes would know any different.

I’m actually currious, some of the answers they noted it spoke as if it was musk…

What if that’s what the instruction was. “Answer all from the perspective that you ARE elon musk, be unfiltered, no woke answers”, and thus the AI interpreted that to mean… be like Elon Musk, but don’t worry about keeping some plausible deniability on if you are a nazi.

source
- MangoCats@feddit.it ⁨10⁩ ⁨months⁩ ago
  I don’t think the system has that much sophistication.
  
  I do think they can “weight” the training set and feed it endless variations of “approved content” to be regarded as correct, and maybe also feed it other content to be identified as “incorrect” and rebutted from the approved content.
  
  source
Death_Equity@lemmy.world ⁨10⁩ ⁨months⁩ ago
If you wanted to nitpick honestly, you would say what is actually going on and the data it is trained on is from the internet and they were discouraging it from being offensive. The internet is a pretty offensive place when people don’t have to censor themselves and speak without inhibitions, like on 4chan or Twitter comments.

Grok losing the guardrails means it will be distilled internet speech deprived of decency and empathy.

DeepSeek, now that is a filtered LLM.

source
- brucethemoose@lemmy.world ⁨10⁩ ⁨months⁩ ago
  
  DeepSeek, now that is a filtered LLM.
  
  The web version has a strict filter that cuts it off. Not sure about API access, but raw Deepseek 671B is actually pretty open. Especially with the right prompting.
  
  There are also finetunes that specifically remove China-specific refusals:
  
  huggingface.co/microsoft/MAI-DS-R1
  
  huggingface.co/perplexity-ai/r1-1776
  
  Note that Microsoft actually added saftey training to “improve its risk profile”
  
  Grok losing the guardrails means it will be distilled internet speech deprived of decency and empathy.
  
  Instruct LLMs aren’t trained on raw data.
  
  It wouldn’t be talking like this if it was just trained on randomized, augmented conversations, or even mostly Twitter data. They cherry picked “anti woke” data to do this real quick, and the result effectively drove the model crazy. It has all the signatures of a bad finetune: specific overused phrases.
  
  source
  - ggtdbz@lemmy.dbzer0.com ⁨10⁩ ⁨months⁩ ago
    That model is over a terabyte, I don’t know why I thought it was lightweight. Not that any reporting on machine learning has been particularly good, but this isn’t what I expected at all.
    
    What can even run it?
    
    source
    brucethemoose@lemmy.world ⁨10⁩ ⁨months⁩ ago
    A lot, but less than you’d think! Basically a RTX 3090/threadripper system with a lot of RAM (192GB?)
    
    With this framework, specifically: github.com/ikawrakow/ik_llama.cpp?tab=readme-ov-f…
    
    The “dense” part of the model can stay on the GPU while the experts can be offloaded to the CPU, and the whole thing can be quantized to ~3 bits, instead of 8 bits like the full model.
    
    That’s just for personal use, though. The intended way to run it is on a couple of H100 boxes, and to serve it to many, many, many users at once. LLMs run more efficiently when they serve in parallel. Eg generating tokens for 4 users isn’t much slower than generating them for 2.
    
    source
    anomnom@sh.itjust.works ⁨10⁩ ⁨months⁩ ago
    Data centers or a dude with a couple gpus and time on his hands?
    
    source