I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.

These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):

  • Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
  • If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?

If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???


edit

Thanks to all who answered.

I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.

Asking to get arguments explained, I got more arguments instead.

    • Wildmimic@anarchist.nexus
      link
      fedilink
      English
      arrow-up
      17
      arrow-down
      1
      ·
      edit-2
      4 hours ago

      You can also train your own local models with license free material if you wish! I think one of the easiest ways to get into that is by using software like unsloth (that’s the one i am using), an open source no-code tool which can be both used to train models on whatever data you wish and to run models either locally or using an inference provider.

      Quick example for something like that which is also not dependent on copyrighted material is RAG, where you can provide the 400-page manual for something and then can chat with “the document” to get explanations, ask quick questions without searching for possibly multiple occurrences of a specific term and similar stuff. I love this for technical documents like mainboard manuals!

    • masterspace@lemmy.ca
      link
      fedilink
      English
      arrow-up
      12
      arrow-down
      3
      ·
      1 day ago

      What’s not ethical is a copyright system thag enforces artificial scarcity where there is no need for it.

      Piracy is not stealing, and is not inherently unethical.

      • dudeface@lemmy.world
        link
        fedilink
        arrow-up
        4
        ·
        1 day ago

        Everything is copyrighted regardless of it being owned by a corporation or blogger

        I don’t partially care about the former

        • partial_accumen@lemmy.world
          link
          fedilink
          arrow-up
          6
          ·
          1 day ago

          The absolutely most generous copyright protection length is in Mexico which is 100 years plus the life of the author when created. So lets generously say a total of 160 years. Most of the rest of the world is 120 year max. Anything outside of that is in the Public Domain. Many things had a much shorter path to the Public Domain falling into it in as little as 20 years.

          So lots and lots of stuff is not copyrighted.

    • Xaphanos@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      1 day ago

      Some don’t speak any “language” they are trained to “speak” and “think” in terms of election orbitals and bonding energy. They are used in pharma and materials science to work on intractable problems like superconductivity and meds for Parkinson’s.

    • Pamasich@kbin.earth
      link
      fedilink
      arrow-up
      3
      ·
      1 day ago

      There is Apertus which at least claims to only use permissively licensed sources and respect robots.txt opt-outs. They outline their methodology and source datasets in this document. Though I haven’t checked the actual sources myself.

    • partial_accumen@lemmy.world
      link
      fedilink
      arrow-up
      5
      arrow-down
      1
      ·
      1 day ago

      As an exercise so I could learn the technology, I trained a model exclusively on the Public Domain works of L Frank Baum, specifically the “Wonderful Wizard of Oz” series (did you know there are 14 books in the series just by Baum?!). The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.

      Again, I didn’t do this to produce a useful tool, but rather an exercise to learn how to build models.

      They are all trained on copyrighted material without permission, no LLM is ethical

      However, this means I can refute your statement because 100% of the input data is public domain novels. Also, I trained it on my own hardware powered 100% by solar power.

      • MonkeMischief@lemmy.today
        link
        fedilink
        arrow-up
        2
        ·
        edit-2
        4 hours ago

        The model created is not useful as a tool, and more often than not just produces English gibberish, but every now and then it can produce original coherent statements.

        So we created multitudes of digital byte-sized monkeys…and gave them m/billions of typewriters. . . :p

        • partial_accumen@lemmy.world
          link
          fedilink
          arrow-up
          2
          ·
          edit-2
          52 minutes ago

          So we created multitudes of digital byte-sized monkeys…and gave them m/billions of typewriters. . . :p

          You’re not far off from describing primative LLMs. In fact, the recent innovation in the last 5 years was the application of a specific technique published in a paper called “Attention Is All You Need” that the secret sauce is “attention” added via a transformer layer. This concept is what made GPT, Claude, etc the life-like responses they have today.

      • Sergio@piefed.social
        link
        fedilink
        English
        arrow-up
        3
        ·
        1 day ago

        Cool that you did that. However, “Large” as in “Large Language Model” generally refers to the dataset size, and 14 books is not Large.

    • feral_sh@lemmy.today
      link
      fedilink
      arrow-up
      19
      arrow-down
      16
      ·
      1 day ago

      Human brains are all trained on copyrighted material without permission, no human brain is ethical