I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
edit
Thanks to all who answered.
I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.
Asking to get arguments explained, I got more arguments instead.


As an exercise so I could learn the technology, I trained a model exclusively on the Public Domain works of L Frank Baum, specifically the “Wonderful Wizard of Oz” series (did you know there are 14 books in the series just by Baum?!). The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.
Again, I didn’t do this to produce a useful tool, but rather an exercise to learn how to build models.
However, this means I can refute your statement because 100% of the input data is public domain novels. Also, I trained it on my own hardware powered 100% by solar power.
So we created multitudes of digital byte-sized monkeys…and gave them m/billions of typewriters. . . :p
You’re not far off from describing primative LLMs. In fact, the recent innovation in the last 5 years was the application of a specific technique published in a paper called “Attention Is All You Need” that the secret sauce is “attention” added via a transformer layer. This concept is what made GPT, Claude, etc the life-like responses they have today.
Cool that you did that. However, “Large” as in “Large Language Model” generally refers to the dataset size, and 14 books is not Large.