I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
edit
Thanks to all who answered.
I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.
Asking to get arguments explained, I got more arguments instead.
Ok seeing your edit, some basic explanations:
A Token is a small unit used by LLMs. Mostly not a full word, more of a word fragment of a few characters.
A Dataset is an accumulation of different data (mostly real-life data to avoid problems) which is used to train a Model.
A Model is basically a huge file, full of Tokens that are interconnected with each other.
Open Weights means that the numbers which describe the relationship between Tokens (this is all basically a ton of matrix multiplication, which is the reason this is using so much power) are available.
If you want to really understand, 3Blue1Brown has a complete series about neural networks!
To answer a lesser answered sub question of a question: what is a weight?
Imagine you have an output, its either a or b. You have an input, whatever it is. And then there is a weight in the middle. The weight determines wether the output will be a or b.
In math it would be a function input*weight=output.
ANN (Artificial Neural Networks, just afancy name) is a chain of a whole bunch of weights. There are layers and each layer generates an output from a bunch of weights and the next layer uses it as inputs. The name comes from the visualization, which looks like neurons firing to other neurons (from biology, things that make your brain work).
Oh and the connection between LLMs, Tokens and ANNs.
ANN = LLMs, LLMs are a subset of ANNs that have the purpose of doing anything language. Most times you’d want maybe to predict a single thing (e.g. a temperature, a color, a letter). The specific output depends on the usecase. LLMs simply output characters and predicts one character at a time whats the most likely output.
To make sense of this nonesens you’d need to output and input things that are larger than a single character. These are tokens. I don’t know much about that, but basically it is the whole reason why LLMs took off and they got invented by some research done by google in 2017 (idk the title of the paper anymore, maybe something along the lines of “a new way of thinking”). Funnily enough google ignored the paper and some guys at google got mad and founded their own AI company cough Anthropic cough
Anyways tokens let the AI-Model predict 4 characters or so at a time, which enables the LLM to predict actual words.
I am not sure exactally how it works from there, maybe it gives the whole in- and output so far back into the model as a new input to generate the next token? Sounds kinda inefficiant, but the whole thing is just fucking inefficient and useless. AI has usecases but LLMs are billionare madness in terms of calculation power needed.
I’ll go one by one… although I do have to say, I can’t answer the main question of whether it can be 100% ethical
Open-weight model
It’s actually quite literal. LLMs may look like magic, but under the hood are basically very, very big mathematical/statistical functions that transform one type of input (for example, “User: hey ChatGPT what is an open-weight model?”) to an output (example, “ChatGPT: To understand what an open model is…”)… To simplify, these models function by piecing together a ridiculous amount of tiny, nonlinear mathematical functions, apply a weight to each of these tiny functions, and somehow get a coherent result out of it. The theory behind this has been known in math research for decades before OpenAI.
Because LLMs are mathematical models, if you know the weights of those tiny functions and how you chain theme together, you can run an exact replica of it. Open-weight models are just that, models where the entire weight and model architecture are released with open, permissive licenses. Many of these end up being modified by other users/hobbyists, and I do believe this sharing/modifying part of AI/LLM is quite true to FOSS philosophy.
As others have mentioned though, “open-weight” has no bearing on how the model is trained. It just means that the final product is free and open
Is it really feasible to run it 100% locally?
Theoretically yes, practically yes but depends on the model.
- For the really small models designed for mobile phones, absolutely. For example, LFM-2.5 and Ling-3.0-tiny run on just about any potato. Granted these models are… not great. They are probably functional as “search engine wrappers” though, methinks that’s what they were built for.
- For the intermediate models, yes but with an asterisk. They require at least a decent GPU, so either you have a gaming rig, happened to have snagged one of the AI computers, or try to kneecap the model in some way (limit the amount it generates, make the mathematical operations less precise a.k.a. quantization, be okay with it running for a long time, …). I have a local setup for Qwen3.8-27B with 4-bit quantization, but it runs at 2 tokens/second on my laptop (which has a ton of RAM but not a very good GPU)…
- For really large models, you can technically run them, but they are realistically designed for organizations with privacy in mind: hospitals, private companies that handle sensitive data, etc…
The software doesn’t come from nowhere
I mean IMO that’s the heart of the whole debate. Making a modern LLM requires massive amounts of data and computation… both of these do have to come from somewhere. And my understanding is that, current AI/LLM research didn’t create that much technological breakthrough, so building a better model often does indeed rely on either using more data or bigger architecture (which takes longer to compute)… which then leads to a truckload of ethics issues here and there
For private companies, we know that OpenAI is blatantly hoarding data from just about everywhere, see the recent Mathematics Millennium Prize scandal. A lot of the open-weight models are trained by Chinese training houses, and Chinese tech companies are not known for respecting intellectual property… and China uses lots of coal energy still.
Maybe Mistral is more ethical, given that 1) they are based in France and care about GDPR and data regulations a lot more, and 2) France uses lots of nuclear energy, but I can’t say for sure. An European company operating an AI/LLM training house from Iceland using their excess geothermal energy is probably the most ethical we can get on this
In theory (and take this with a whole jug of salt) someone can run an off-grid farm and train a small LLM all on their own, which would be very ethical… but I doubt any serious organizations are doing it this way. Most ppl just want better performance and/or more optimized models for now
For a complete FOSS purist with an environmentally conscious angle, I don’t know if there is an AI/LLM that can be considered 100% ethical… but then again Gentoo is a thing. So yeah, I don’t know
Is it really feasible to run it 100% locally?
Yes. My mid 7-year-old rig featuring Ryzen 5 1600X, 16 GB DDR4 RAM, and GTX 1060 6gb can run fairly intelligent models that are actually practical to use.
The software doesn’t come from nowhere
Correct. Training the model is the most resource-intensive and planet-churning part.
In my book, not paying and supporting them financially - > ethical
My issue is when you pay for a subscription, when you buy a gpu specifically for ai, when you buy a model, etc. Because then you are funding the slop machine.
There are ethically-sourced datasets used to train LLMs that can be powered by renewable energy and cooled by closed water systems. Maybe the question becomes “sure, but are those good for anything?” and the answer to that is “probably not right now”.
If you can play modern games, you can probably run a local LLM. The more VRAM the better it’ll go.
For ethics, there’s two issues. The first is power consumption. Running a model locally isn’t consuming any more power than running a game that fully pushes your system. It’s not really a concern.
The second issue is content generation. It’s all created through the theft of content others have made. There’s no getting around that. As far as I’m aware there are no ethical models that specifically have not stolen content. If you’re just using it to chat with or whatever, I think it’s fine. You probably aren’t taking work from anyone, at least to any degree more significant than piracy. If you’re generating stuff to be distributed, then you’re distributing stolen work. If that’s ethical or not is up to you to decide.
For ethics, there’s two issues.
No, there’s seven main issues.
I feel like this question is asked in bad faith if you are able to list seven “main” ethical issues here. I don’t know if I could list seven different issues that aren’t just breaking down power consumption and content theft into smaller pieces, and I’m pretty anti-AI. I guess you could include hardware consumption too. Financing maybe, but that doesn’t play a part in local models generally. Care to explain?
Then you misunderstood the question. Re-read the title?
There’s a lot of info that you need to know to explore this space, so I’ll take my own shot at answering. Let me know if anything needs further explaining!
An “open weight” model is an AI model where you can download the data needed to run the model on your own hardware for free. Contrast this with proprietary models-as-a-service like ChatGPT and Claude where you have no access to the data needed to run the model yourself – you can only use it through the services provided, usually for a fee, and which can be taken away from you or changed at any time with no recourse.
The mapping to traditional open source terms does not work well since what you get is a binary artifact.
Those artifacts are released with a license – and many of the models are licensed permissively (e.g. MIT or Apache license terms). You can take those weights, modify them, and then release them as new models – and people do actually do this in practice!
Is it really feasible to run it 100% locally?
Yes. I run models on my own computers and have tried a number of configurations to figure out what works well. The Qwen family of open weight models (from Alibaba) are the ones I’ve found most useful so far. Gemma4 models (from Google) are also useful.
I prefer models that have been modified by the community to remove corporate censorship – i.e. stripping that “As a large language model…” cover-your-ass crap and evasiveness on topics like Tiananmen Square. If that means the model is technically capable of telling me to go kill myself too, so be it; I’ve spent 25+ years dealing with assholes on the internet and can handle abuse from a stupid robot if I have to. (In practice though, they’re usually pretty nice still unless I deliberately tell them to act like an asshole – and then Qwen, at least, starts to sound like a snarky redditor; it’s quite funny most of the time, actually.)
If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
Models do require training to create, yes. It’s generally not clear where the training physically happened IRL – so, yes, some of them probably used power from gas turbines, but others may be drawing power from the Three Gorges Dam in China or solar plants or nuclear plants or whatever else is hooked up to the electric grid where the training happened. Most of them are also not very open about the data sets they were trained on. (There are exceptions to this though!) The Chinese models in particular are almost certainly trained heavily on logs extracted from Western models in addition to using whatever other data they could get ahold of. Whether you think that’s ethical or not is a matter of perspective; how do you feel about Robin Hood?
Once a model has been trained though, it can run on a normal GPU. The power requirements to run an LLM are basically the same as running a video game, or, equivalently, about the same as turning on a few incandescent lightbulbs. (The iGPU in one of my systems uses 100W; the discrete GPU in another system I’ve tried uses 215W under load with appropriate tuning – or 300W if you run it naively.)
If you want to run a model yourself, I recommend using llama.cpp – there are instructions on how to get started with it here: https://llama.app/
These are the models I’ve found most useful:
- Qwen3.8-27B (official version): https://huggingface.co/Qwen/Qwen3.8-27B
- Qwen3.8-27B (uncensored): https://huggingface.co/MuXodious/Qwen3.8-27B-absolute-heresy
- Qwen3.6-35B-A3B (official version): https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- Qwen3.6-35B-A3B (uncensored – ⚠️ mildly NSFW graphics on page): https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic
- Gemma4 (official collection – many variants in different sizes): https://huggingface.co/collections/google/gemma-4
If you have an iGPU only, I recommend using one of the so called “Mixture of Experts” (MoE) releases. These are typically named like 35B-A3B or similar; the first number indicates the total number of weights (35 billion) and the second indicates how many are “active” (i.e. actually used during computation) at one time while the model is running (3 billion in the example). These models need less computation to run and stay fast on weaker GPUs. Qwen3.6-35B-A3B is very good in this space and was my go-to model for a long time.
If you have a discrete GPU and enough VRAM, I recommend using a dense model (i.e. one that activates all its weights while answering) like Qwen3.8-27B.
It’s worth noting that people don’t usually use the full quality weights (which are typically ~2 bytes per weight); they use a “quantized” version – compressed in a lossy fashion like a JPEG. Going down to 4-bits (half a byte) on average per weight is about as low as most people like to go – you will see this indicated in names like Q4_K_M. (Quantized to ~4 bits with the K quantizaation scheme, medium variant.) Usually a bigger number is better in the sense of “closer to the original quality” – at the cost of needing more RAM.
Full quality weights are often found as safetensor files on HuggingFace. Quantized weights intended for use with llama.cpp are usually in GGUF file format.
Does that help?
Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
Very useful 35B models need (more or less) 8GB of VRAM and 32GB+ CPU RAM to be usable locally. 16GB RAM might work on a lean system with an RTX Nvidia card and an exl3 quantized model.
I know we are in a RAM apocalypse, but pre-apocalypse, that’s a quite reasonable requirement, IMO.
Personally, I run Deepseek V4 07-31 Flash at 19 tokens/second on a desktop with a single RTX 3090 and 128GB CPU RAM, and that’s an extremely capable model. Again, that’s expensive these days, but pre-ram apocalypse, that is not an unreasonable workstation/homelab.
If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters…
Most open weights LLMs are Chinese. And they are:
-
Trained on pennies. They have to be, as they simply do not have a sea of GPUs like Big Tech. Training costs for their large models are in the millions or tens of millions; a single steel forge has used more energy than all of those training runs combined.
-
China relies more on renewables, and I believe the datacenters aren’t so hastily constructed with gas turbines in the middle of cities.
-
And as of now, they are transitioning away from Nvidia GPUs. Some labs already use Huawei accelerators.
…and stolen IP and stolen personal data?
Yep.
This is a huge caveat.
You can avoid this. Nvidia Nemotron models, for example, are trained on completely open datasets you can download and inspect yourself: https://huggingface.co/nvidia
They are very good, but just behind state of the art.
But in practice, the SOTA models most run use private datasets. Lord knows where the Chinese get it from, but given some common quirks between models, at least some data sources are shared (and possibly government provided?)
…However.
I would argue providing the result of the training as Apache licensed weights counts as “fair use,” in the same way non commercial fan works do.
They aren’t making a dime off releasing those weights. I’m not trying to sell anyone anything when I use them. Where is the IP theft if money isn’t changing hands?
Now, the Chinese LLM services they charge for? I have no excuse for that. Once money is on the table, it is definitely IP theft.
completely open datasets
Not completely – it is mostly open, but they use a dozen or so private datasets for things like training on global regulations, minesweeper (for some reason), etc. To their credit, they do indicate this on the model cards, but it’s not entirely clear what is in those datasets either.
(I fell for that bit of marketing myself awhile back.)
Where is the IP theft if money isn’t changing hands?
Money not changing hands is pretty much the definition of theft.
Unlike the definition of heft which is a photo of OP’s mom.
What I’m saying is it’s akin to writing a fanfic or making fanart of your favorite franchise. Or getting inspired by a painting you see, and making something similar yourself.
Do that for your personal enjoyment? That’s fair use, under the law.
But the moment you start trying to sell it is when you get in legal hot water, and when it’s indeed morally problematic.
The scale is different, but I’d argue a similar principle applies: if you use some model trained on public works from a protected IP, but the model and its outputs are not resold, nor profited from, it’s not theft. The point is beyond money not changing hands; there’s no profit being made from the original author’s stuff. They aren’t being taken advantage of any more than someone viewing their public stuff for free, or someone creating derivatives from private work.
But all that is off the table the moment profit and distribution is involved.
-
From that perspective local LLMs sound more like classical piracy. Not ethical, not FOSS, but out of the hands of greedy corporates.
they are also doing distillation from the big models from OpenAI & Claude, so open-weight models with similar capability can be available for free and also reduce the big AI companies’ ability to profit from stolen data.
What are the most successful/useful distillations you can run locally?
It’s the backend stuff I wanna figure out so it can convert plain language requests and notes into personal assistant tasks while being flavorfully bitchy about all of it. Apparently RAG and agent tools have something to do with it.
I’ve heard good things about Qwen.
Right… I meant like which specific file.
Qwen.
Some of the open models claim not to use pirated content. I don’t necessarily believe it, but its being claimed.
But also out of the hands of the small artists/devs they’re pirating.
Its much less targeted piracy than the traditional, and its almost weird to see the rules pirates implicitly respected being broken.
Is it really feasible to run it 100% locally?
Yes. My Macbook Air (M2) released in 2022 can run many publicly available LLM models. The ouput is not as fast as using a large powerful datacenter, but for my local needs, I’m not in a hurry. I get about 17 to 30 tokens per second speed running 100% locally.
If yes to the previous: the software doesn’t come from nowhere
For Mac users, the interface comes from here . For the specific LLM models that is a separate question for each.
and ultimately still relies on gas-turbine-powered datacenters
DCs generating power on-site is a relatively new phenomon because existing grids are at capacity, so the only way to bring new DCs online is locally generating power at that DC, usually using gas turbines or even worse, diesel generators. Most if not all of the publicly available models for running on your own hardware were built before those gas-turnbine-generating DCs were a thing.
The public models people are running now have existed for a number of years are likely made on regular utility grid power which is whatever that nation and region uses.
and stolen IP and stolen personal data, no?
The Llama LLM is made by Meta, so probably yes for that one. Deepseek is from an AI research lab in China. QWEN is from Chinese company Alibaba. We don’t know for sure the inputs that created the Chinese models. US AI companies claim a number of the Chinese models are derived from American LLMs, but I haven’t seen (or looked for) proof of these claims.
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically
Likely they mean because you don’t have to pay a large LLM owner in a rent-seeking model to run LLMs, nor does a person’s use contribute to further development by those companies.
or true to FOSS philosophy, because … ???
With the open-weight models it doesn’t rely on a commercial license to use, and the interfaces can be truly open source.
OK but what does open-weight model mean?
“Weights” in LLMs are the “final answer” numbers from the results of model training, and these are the engine of the LLM model. Lets use chocolate chip cookies as an analogy.
The chocolate chip cookies are produced with ingredients, a recipe, labor effort to produce the dough, and then cooking energy/effort to bake chocolate chip cookies. In a traditional AI company, the company gets the ingredients, they write their own recipe, do all the dough creation, and then the energy/labor for baking, and charge you money to get chocolate chip cookies.
An open-weights AI company got all the ingredients, wrote their own recipe, labor effort to produce the dough, and instead of charging you for the dough, you get as much uncooked dough as you want forever. Your only task is to take the dough and cook it yourself and you’ve got free chocolate chip cookies. You can make as many cookies as you want with your own oven. However, you are not given free raw ingredients, nor are your given the recipe to alter it in a way you might like. You can only get the dough for free.
So open-weight AI models (uncooked dough) are LLMs you can use on your own computers (oven) for free and have LLM output ( chocolate chip cookies), but you don’t get the training data (ingredients) nor the training parameters (recipe) that built the open-weight model.
Thanks, that was properly eli5’d!
It means the weights (the numeric data that makes up the model) are publicly available to download and use for free.
In some cases there are conditions such as, if you run the model as your business, you need to purchase a license. That’s usually the largest, most powerful models, though, not those most people would be able to run at home.
Question, what do you do with it and how is the data that underlies the model harvested?
Question, what do you do with it
I use it for both fun throwaway stuff as well as productivity tasks for learning.
Example of fun throwaway:
Do you remember that episode of Seinfeld where the Kramer starts using Facebook marketplace and starts buying the most worthless items before being robbed when trying to get a too-good-to-be-true sale? No? Because it never happened, but you can plug that premise into an crafted LLM prompt and it will pop out a whole TV script with in-character dialog for each actor as well as use of popular existing sets.
Foreign language learning:
I’m studying a foreign language and want to interact with just the level of vocabulary and grammar I have knowledge of right now at my level for practice. I can prompt the LLM to limit itself to just what I know now and adjust the conversation level so I can practice. If I ever get stuck, I can ask the LLM to explain the grammar usage or vocabulary choice.
and how is the data that underlies the model harvested?
I’m not sure what you’re asking here. Are you asking, for example, how the Deepseek model was trained? If so, I answered that above. If not, can you rephrase your question?
You run a program like llama.cpp that can use the weights to run the model, which you can then use for whatever you’d use a model for: figuring out tech stuff, coding, etc. It’s a bit involved, but there are tools that make it easy to get started. LM Studio for instance.
In some cases, the lab that created the model publishes the dataset that it was trained on, and those are usually made of publicly available data. In most cases, though, the labs don’t give details, but the answer likely involves siphoning every web page they could find.
They are all trained on copyrighted material without permission, no LLM is ethical
You can also train your own local models with license free material if you wish! I think one of the easiest ways to get into that is by using software like unsloth (that’s the one i am using), an open source no-code tool which can be both used to train models on whatever data you wish and to run models either locally or using an inference provider.
Quick example for something like that which is also not dependent on copyrighted material is RAG, where you can provide the 400-page manual for something and then can chat with “the document” to get explanations, ask quick questions without searching for possibly multiple occurrences of a specific term and similar stuff. I love this for technical documents like mainboard manuals!
Doesn’t RAG require a pretrained model still? Presumably on copyright material?
It does not per se need a model using copyrighted stuff. You can get by using a model without that - basic language skills and technical jargon can easily be trained with open material.
In theory that makes sense, but does this actually exist?
It does if you want to go that route. https://www.kaggle.com/datasets is a good starting point.
Thanks for the info!
What’s not ethical is a copyright system thag enforces artificial scarcity where there is no need for it.
Piracy is not stealing, and is not inherently unethical.
Everything is copyrighted regardless of it being owned by a corporation or blogger
I don’t partially care about the former
The absolutely most generous copyright protection length is in Mexico which is 100 years plus the life of the author when created. So lets generously say a total of 160 years. Most of the rest of the world is 120 year max. Anything outside of that is in the Public Domain. Many things had a much shorter path to the Public Domain falling into it in as little as 20 years.
So lots and lots of stuff is not copyrighted.
Some don’t speak any “language” they are trained to “speak” and “think” in terms of election orbitals and bonding energy. They are used in pharma and materials science to work on intractable problems like superconductivity and meds for Parkinson’s.
that’s not really an LLM then, is it?
It is, cause it uses the same architecture, https://www.geeksforgeeks.org/artificial-intelligence/exploring-the-technical-architecture-behind-large-language-models/
an example: https://github.com/bowang-lab/scGPT “talks” in RNA and is incredibly accurate and able to do predictions
But it is AI.
true, but that’s hardly relevant as a reply to s comment which is clearly referring to LLMs and not machine learning in general
There is Apertus which at least claims to only use permissively licensed sources and respect robots.txt opt-outs. They outline their methodology and source datasets in this document. Though I haven’t checked the actual sources myself.
As an exercise so I could learn the technology, I trained a model exclusively on the Public Domain works of L Frank Baum, specifically the “Wonderful Wizard of Oz” series (did you know there are 14 books in the series just by Baum?!). The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.
Again, I didn’t do this to produce a useful tool, but rather an exercise to learn how to build models.
They are all trained on copyrighted material without permission, no LLM is ethical
However, this means I can refute your statement because 100% of the input data is public domain novels. Also, I trained it on my own hardware powered 100% by solar power.
The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.
So we created multitudes of digital byte-sized monkeys…and gave them m/billions of typewriters. . . :p
So we created multitudes of digital byte-sized monkeys…and gave them m/billions of typewriters. . . :p
You’re not far off from describing primative LLMs. In fact, the recent innovation in the last 5 years was the application of a specific technique published in a paper called “Attention Is All You Need” that the secret sauce is “attention” added via a transformer layer. This concept is what made GPT, Claude, etc the life-like responses they have today.
Cool that you did that. However, “Large” as in “Large Language Model” generally refers to the dataset size, and 14 books is not Large.
do you consider “piracy” like zlibrary and annas archive unethical?
If a business is using them, yes.
i agree.
in that case, wouldn’t it be similar ethically if people use these models on their own machine for personal use?
Human brains are all trained on copyrighted material without permission, no human brain is ethical
You can run them locally, yes. There are models that can even run on phones, but usecase is limited. But it can only be considered ethical, if the training data used is listed or ethically sourced IMO.
AI bros on Lemmy will disagree with me, but most open weight models are still trained unethically i.e, theft. Most proponents of LLMs (who I talked to on bsky), who say local models are ethical, don’t fucking use it. They’re larping on socials about how awesome it is, but none of the ones I talked to are using it in their projects. They mess around, realise it is not as good as the “unethical” options, go right back to Claude
Open weight models Qwen, deepseek, mistral, and the Ollama stuff etc are unethical in normal people’s eyes, but “ethical” enough for AI bros.
From what I searched, there are very few that can be considered ethical - Olmo, Apertus, Starcoder(?). But idk anyone who uses these. My friend at IBM said they used Apertus, but it was nowhere near good as ChatGPT, so they no longer use Apertus now. And these models require minimum 6-8 GB VRAM for their lowest parameter model iirc.
Even the open-weight model bros are lobbying to redefine what ‘open-source AI’ means. That should give you a fair idea about people behind open-weight as well
I would unironically argue that a model primarily trained through distillation of closed frontier models, and then released open-weight with an open-source architecture, becomes “ethical” again.
Something something Robin Hood
Rob the poor’s money from the rich and keep it for your community?
Sounds more like feudal warfare than anything, I can’t see any harm to artists being reduced at all
Borderline strawman there but I’ll bite.
Open weight models trained unethically are unethical. Closed weight models trained unethically and then sold back to you for a profit from gas-powered datacenters funded through Ponzi schemes are substantially more unethical.
From there it’s a harm reduction calculus. No, those are never pleasant.
So do you let the closed weight labs conquer the field unopposed just so you can feel better about yourself? That’s a valid stance, FWIW, and it’s also super easy and convenient because you don’t have to do anything. It especially makes sense if you believe it’s still possible that AI will just go away on its own. I don’t, myself, not anymore, so I encourage the use of open weight models, however grudging, so people don’t give money to the closed labs and in the worst case aren’t eventually stuck with the maximally unethical options. And we’ve not even touched on the nightmare labor replacement scenarios that seem every day less unlikely. Am I right? I have no clue. Like you, I’m just trying to make the best choices I can in a world that’s gone to shit. I’d recommend dropping holier-than-thou attitude either way, though, because it doesn’t help our side. Man.
Just as I suspected…
Thanks for taking the time.
there are very few that can be considered ethical - Olmo, Apertus, Starcoder
This is software meant to be run always and completely locally?
Sorry to whine, but so far nobody has eli5’d what “open-weighted” means, or “model” at that… please?
This is software meant to be run always and completely locally?
That’s what is claimed, I haven’t run them locally since I don’t have a good system.
To be honest, I’m not sure if I can eli5 weights and models, but I’ll try. Think of a model like the base - for example, OpenAI has different models like Astra, Sol, etc. These are different models, like different versions of a software or operating system like macOS, but for AI stuff. Like one would download a software, you download a model to perform tasks.
Weights are vales that can influence inputs of these models to get a desired/better result. What most of these models are doing is mostly predicting what might be the next appropriate text/data to the question you asked. When you ask these AI models what 2+2 is, it is not performing a math operation like a normal program, it is looking at its training data to see what the closest option might be. It is doing pattern matching.
These AI models inside can be thought of like an interconnected network, like neurons in our body, that keep passing information to the next neuron and to the brain to make a decision. (Before understanding LLMs it would help to understand Neural Networks first). These AI networks need weights and biases. These networks perform calculations and weights are used to determine how much importance/weight each input can have on the output. Bias on the other hand, is used to shift/change the output so the AI model can ‘learn’ to pattern match better.
What open-weight models, do is they make the model available for download along with the weights. No information is given on training data. Like with ads, ones with most data emerges victorious i.e, has a better model. So these companies do theft, don’t list their training data afraid of getting caught. I forgot which one, but either Deepseek or Qwen (both open-weight) was caught ‘stealing’ from Claude (not open weight). You can probably guess how much these companies value ethics.
I’m not sure if this entire thing goes away, but local models might be the ones left standing when this bubble pops.
Open weight is different to open source. Open Source AI as it stands, the definition requires a model to have entire thing made public - so the weights, biases, training data used, the model. Apertus, Olmo etc are mostly meeting open source AI definition.
If you need to know more, this is what we’d use to refresh our memory before exams :)
I probably might have made mistakes here, English isn’t my first language either. But I hope you get an idea about these terms
If you really need to understand this tech more, I recommend watching ‘AI for Everyone’ course on Coursera from Andrew Ng. It is free to audit, my friends who took his course were hyped (I wasn’t really interested in AI)
Thanks a lot.
I did not know it was possible to influence AI software/models and nudge them in a certain direction.
These AI companies have even more power than I thought for a long time.
I’ll take a crack at it.
I run AI on an old Nvidia p40 purchased online for less than $200. The models I run on it came predominantly from Chinese companies and groups. On balance, they:
-
use much less fossil fuel in producing these LLMs than their Western counterparts
-
produce models that run much better on lower end consumer hardware
As for training data, the inputs used to create these models, it varies greatly but the most popular line, Qwen, Came from the company’s own data from operating such huge networks and systems for so long
I really fail to see how any of that is worse than playing a video game.
None of this takes away from the very real issues around data center build out and Western companies using the systems to scare people and continue an economic bubble. That’s all true and bad. But there’s nuance. Not all AI is created equal
use much less fossil fuel in producing these LLMs than their Western counterparts
Utterly false, given that all the decent Chinese models (especially Qwen) are distilled from Western frontier models.
They literally could not exist without the enormously carbon emissive western models existing first to train them.
Also China has an insanely diversified grid, that utilises almost a global scale of fossil fuels.
So its like green communism at best
distilled from Western frontier models.
What does this mean please (remember the title of the post!) and how do you know?
Basically, it’s really really expensive (from a compute standpoint, and more compute = money and energy) to train a decent LLM just using books and conversations etc. Anthropic, OpenAI, etc spend a truly insane amount of money doing this is in data centers.
Running an LLM like that is also extremely expensive from a compute standpoint, and right now the true frontier models can basically only be run in datacenters.
However, once you have a really good LLM, you can use it to train another LLM. Because it’s basically already refined all its original training data (and is capable of further refining it’d output on the fly), training a new LLM from it is vastly easier and cheaper.
Additionally, you can use a much smaller and more efficient model, so it can run on less powerful hardware.
So all local LLMs that are any good that exist right now, were trained using the outputs of existing super expensive frontier models. They could not exist without the frontier models to train them. This is also how OpenAI and Anthropic create their cheaper models. Mythos/Fable is anthropics current frontier model and they distill it into (use it to train) their cheaper models like Sonnet and Opus.
Wow. as if the original models weren’t rickety enough.
They’re honestly not that rickety. If you’re getting your news from social media (including Lemmy) you may want to actually try them rather than just take the random incidents where things go wrong as broadly representative.
yeah the comment was bad enough I looked at the user. been around for 2 years and no posts or comments till this one. made a note on the account.
-
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
I’m not holding this as a hard stance, just a current analysis of the technology and how it might affect my life in my circumstances. I think we’re well into a world where we need to be deciding if some technologies don’t suit our lives, because there certainly are even more harmful technologies to come.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
In leaving progress to the machines, in letting technology go forward on its own terms and selecting from it, with what seems to us excessive caution, modesty, or restraint, the limited though completely adequate implements of their cultures, is it possible that in thus opting not to move “forward” or not only “forward,” these people did in fact succeed in living in human history, with energy, liberty, and grace?
Always Coming Home - Stone Telling Part 3 by Ursula K. Le Guin
I’ve come round to the idea that I will not use LLMs regardless of how “ethically” they can be sourced because their outputs are harmful in a way that builds up, like a very small dose of poison.
How is LLM output any more harmful than output from a human you don’t know? I would agree with you if one were to simply blindly accept anything an LLM gives you as factual. However, I am skeptical of what humans say too. Some are inaccurate out of carelessness, some out of malice. Critical thinking is the key to protection from both errand LLM output and fallible humans.
Additionally, this thread is about locally running LLMs. One of the most dangerous aspects of most large LLMs is the sycophancy where the LLM will try to tell you what you want to hear, even if it needs to creatively invent things that don’t exist or are not true. If you are not aware, running locally means you have all the controls on the model. You’re not subject to whatever settings a large hyperscaler sets up for you. This means, among other things, you can turn down the “temperature” setting, which is the lever that controls how creative an LLM is. Practically what this means is, if you set it to “0” you are allowing it no creative action. If you ask it a question it doesn’t have actual training on, it will tell you effectively “I don’t know” instead of making something up just to have an answer.
It might seem like a romantic or artist’s approach, but with so much spiralling out of our control I’m trying to make the active choice to be more human.
The older I get the more disappointed I get in humanities large group decisions and actions.










