hendrik@palaver.p3x.de 16 hours ago
Yes. “open-weights” means they publish the AI model file to the public and people can download and run it themselves.
And there’s very different AI models out there. Some are smaller, some bigger… Bigger usually means more “intelligent” and you also need more resources to run them. On average, if you own something like a beefy gaming computer, you can do a reasonably “clever” AI at reasonable speed. Something like ChatGPT needs a datacenter, and at the other end we also have some small models which run on some smartphones or average laptops. They surely won’t be able to do the same things ChatGPT does… And if your laptop is slow, it’ll output text very slowly. But maybe you’re okay with less “intelligent” AI.
And no, they don’t come from nowhere. They’re mainly made by the AI companies. They’re trained the same way all large language models are trained with the same problematic procedure. Oftentimes they publish some more information like a scientific paper alongside. There might be information inside about the energy used / carbon footprint of the training steps. Sometimes/rarely they also give information on what data they used.
And internet people have all sorts of weird opinions on (FOSS) philosophy and ethics. I think we’d need to be more specific than that. Dirty gas turbines and allowing companies to be exempt from copyright while other people get sued for copying a single movie isn’t really ethical by any means.
Thanks.
So “they” are the big corpos involved in AI? I have my hard time wrapping my head around this explanation (although somebody else said the same so it must be true /s). What’s open-weight supposed to mean here? And, usually the thing you get for free is inferior to the paid vrsion - how does that work?
hendrik@palaver.p3x.de 6 hours ago
It’s big US companies like Meta (who own Facebook, WhatsApp etc), Google, Chinese companies like DeepSeek, Alibaba… Some other various startups and AI companies or spinoffs. And the occasional university or research institute. There’s a select few European ones. And Nvidia (who sell the graphics cards and hardware), they do research and publish stuff as well.
I think “open-weights” has been coined because the companies will mislead people and advertise with “open-source”, as in open-source software (like Linux, Firefox etc). But their models rarely include the sources (so to speak). That’d be the training recipe and all the data that went in. They can’t publish the training data though, because they regularly get sued by the people they stole the books from. So… they only publish the resulting AI model.
It’s a bit like an executable file on a computer or a purchased game which you can run at your own terms, on your computer, without any online stuff or anticheat attached. You just don’t own any of the development resources.
That means we don’t know what went in, we can’t recreate it, and we might not be able to learn a lot. We can however run and use the model. (We can even modify it to some degree.)
Of course that’s too easy. It’s not really an executable file. Those AI models are neural networks. All the information and “knowledge” is stored in the parameters of that network. That’d be the weights. Lots of numbers for all the nodes and edges which make up the graph/network.
And not sure about the inferior/superior models… I mean it’s hard to compete with OpenAI or Anthropic and the huge pile of money they have. They hired a lot of talent… And their results will be a trade secret, only offered as a service… But then there’s this whole AI war going on between the US and China. And all the companies against each other. They’re constantly trying to outcompete each other. And usually it won’t take long until someone claims they made an open-weights model as good as (or better than) the current version of ChatGPT.