overview for ClamDrinker

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 1 points 46 minutes ago* (last edited 38 minutes ago)

That was your implied argument regardless of intent.

I decide what my argument is, thank you very much. Your interpretation of it is outside of my control, and while I might try to avoid it from going astray, I cannot stop it from doing so, that's on you.

Completely wrong, which invalidates the point you want to make. “Analysis” and “as is” have no place in the definition of copyright infringement. A derivative work can be very different from the original material, and how you created the derivative work, including whether you performed whatever you think “analysis” means, is generally irrelevant.

I wasn't giving a definition of copyright infringement, since that depends on the jurisdiction, and since you and I aren't in the same one most likely, that's nothing I would argue for to begin with. In the most basic form of plagiarism, people do so to avoid doing the effort of transformation. More complex forms of plagiarism might involve some transformation, but still try to capture the expression of the original, instead of the ideas. Analysis is definitely relevant, since to create a work that does not infringe on copyright, you generally can take ideas from a copyrighted work, but not the expression of those ideas. If a new work is based on just those ideas (and preferably mixes it with new ideas), it generally doesn't infringe on copyright. It's why there are so many copycat products of everything you can think of, that aren't copyright infringing.

No it detects patterns. You already said it correctly above. And the problem is that some patterns can be copyrighted. That’s exactly the problem highlighted here and here. For copyright law, it doesn’t matter if, for example, that particular image of Mario is copied verbatim from the training data.

While depending on your definition Mario could be a sufficiently complex pattern, that's not the definition I'm using. Mario isn't a pattern, it's an expression of multiple patterns. Patterns like "an italian man", "a big moustache", "a red rounded hat with the letter 'M' in a white circle", "overalls". You can use any of those patterns in a new non-infringing work, Nintendo has no copyright on any of those patterns. But bring them all together in one place again without adding new patterns, and you will have infringed on the expression of Mario. If you give many images of Mario to the AI it might be able to understand that those patterns together are some sort of "Mario-ness" pattern, but it can still separate them from each other since you aren't just showing it Mario, but also other images that have these same patterns in different expressions.

Mario's likeness isn't in the model, but it's patterns are. And if an unethical user of the AI wants to prompt it for those specific patterns to be surprised they get Mario, or something close enough to be substantially similar, that's on them, and it will be infringing just like drawing and selling a copy of Mario without Nintendo's approval is now.

The character likeness, which is encoded in the model because it is in fact a discernible pattern, is an infringement.

You have absolutely no legal basis to claim they are infringement, as these things simply have not been settled in court. You can be of the opinion that they are infringement, but your opinion isn't the same as law. The articles you showed are also simply reporting and speculating on the lawsuits that are pending.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 1 points 2 hours ago* (last edited 2 hours ago) (1 children)

That's a very short example, but it is a new arrangement of the existing information. It's not a new valuable arrangement of information, but new nonetheless. And yes, rearrangement is transformation. It's very low entropy transformation, but transformation nonetheless. Collages and summaries are in fact, a thing that humans make too.

Unless you mean "new" as in, something nobody's ever written before, in which case not even you can create new information, since pretty much everything you will ever say or write down can be broken down into pieces that have been spoken or written before, which is not exactly a useful distinction.

There’s no transformation, it’s not capable of transformation, it’s just a very complicated text jumbler that’s supposed to jumble text so that the output is readable by humans.

Saying it doesn't make it true, especially when you follow it up with a self-debunk by saying it transforms the text by jumbling it in specific ways that keep it readable to humans, which requires transformation as like you just demonstrated, randomly swapping words does not make legible text..

You’re taking investment advice from a parrot that had the entirety of reddit investment meme subreddits beamed into its brain.

???

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 1 points 3 hours ago* (last edited 3 hours ago) (2 children)

No, not what I said at all. If you're trying to say I'm making this argument I'd urge you (ironically) to actually analyze what I said rather than putting words in my mouth ;) (Or just, you know, ask me to clarify)

Copyright infringement (or plagiarism) in it's simplest form, as in just taking the material as is, is devoid of any analysis. The point is to avoid having to do that analysis and just get right to the end result that has value.

But that's not what AI technology does. None of the material used to train it ends up in the model. It looks at the training data and extracts patterns. For text, that is the sentence structure, the likelihood of words being followed by another, the paragraph/line length, the relationship between words when used together, and more. It can do all of this without even 'knowing' what these things are, because they are simply patterns that show up in large amounts of data, and machine learning as a technology is made to be able to detect and extract those patterns. That detection is synonymous with how humans do analysis. What it detects are empirical, factual observations about the material it is shown, which cannot be copyrighted.

The resulting data when fed back to the AI can be used to have it extrapolate on incomplete data, which it could not do without such analysis. You can see this quite easily by asking an AI to refer to you by a specific name, or talk in a specific manner, such as a pirate. It 'understands' that certain words are placeholders for names, and that text can be 'pirateitfied' by adding filler words or pre/suffixing other words. It could not do so without analysis, unless that exact text was already in the data to begin with, which is doubtful.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 1 points 4 hours ago* (last edited 4 hours ago)

Yes, this is my exact issue with some framing of AI. Creative people love their influences to the point you can ask them and they will point to parts that they reference or nudged to an influence they partially credit to getting to that result. It's also extremely normal that when you make something new, you brainstorm and analyze any kind of material (copyrighted or not) you can find that gives the same feelings you desire to create. As is ironically said to give comfort to starting creatives that it's okay to be inspired by others: "Good artists copy, great artists steal."

And often people very anti AI don't see an issue with this, yet it is in essence the same as the AI does, which is to detach the work from the ideas it was built on, and then re-using those ideas. And just like anyone who has the ability to create has the ability to plagiarize or infringe, so does the AI. As human users of AI we must be the ones to ethically guide it away from that (Since it can't do that itself), just like you would not copy-paste your influences into a new human made work.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 2 points 4 hours ago

For OpenAI, I really wouldn't be surprised if that happened to be the case, considering they still call themselves "OpenAI" despite being the most censored and closed source AI models on the market.

But my comment was more aimed at AI models in general. If you are assuming they indeed used non-publicly posted or gathered material, and did so directly themselves, they would indeed not have a defense to that. Unfortunately, if a second hand provided them the data, and did so under false pretenses, it would likely let them legally off the hook even if they had every ethical obligation to make sure it was publicly available. The second hand that provided it to them would be the one infringing.

If that assumption turns out to be a truth (Maybe through some kind of discovery in the trial), they should burn for that. Until then, even if it's a justified assumption, it's still an assumption, and most likely not true for most models, certainly not those trained recently.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 0 points 4 hours ago (4 children)

They are not “analyzing” the data. They are feeding it into a regurgitating mechanism. There’s a big difference. Their defense is only “good” because AI is being misrepresented and misunderstood.

I really kind of hope you're kidding here. Because this has got to be the most roundabout way of saying they're analyzing the information. Just because you think it does so to regurgitate (which I have yet to see any good evidence for, at least for the larger models), does not change the definition of analyzing. And by doing so you are misrepresenting it and showing you might just have misunderstood it, which is ironic. And doing so does not help the cause of anyone who wishes to reduce the harm from AI, as you are literally giving ammo to people to point to and say you are being irrational about it.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker -1 points 4 hours ago* (last edited 4 hours ago) (3 children)

You say it's not capable of producing anything new, but then give an example of it creating something new. You just changed the goal from "new" to "valid" in the next sentence. Looking at AI for "valid" information is silly, but looking at it for "new" information is not. Humans do this kind of information mixing all the time. It's why fan works are a thing, and why most creative people have influences they credit with being where they are today.

Nobody alive today isn't tainted by the ideas they've consumed in copyrighted works, but we do not bat an eye if you use that in a transformative manner. And AI already does this transformation much better than humans do since it's trained on that much more information, diluting the pool of sources, which effectively means less information from a single source is used.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 1 points 4 hours ago

Which is why the technology itself isn't the issue, but those willing to use it in unethical ways. AI is an invaluable tool to those with limited means, unlike big corporations.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 2 points 16 hours ago* (last edited 16 hours ago) (1 children)

Not 1:1, overfitted images still have considerable differences to their original. If you chose "reproduce" to make that point, that's why OP clarified it wasn't literally copying training data, as the actual data being in the model would be a different story. Because these models are (in simplified form) a bunch of really complex math that produces material, it's a mathematical inevitability that it produces copyrighted material, even for calculations that weren't created due to overfitting. Just like infinite monkeys on infinite typewriters will eventually reproduce every piece of copyrighted text.

But then I would point you to the camera on your phone. If you take a copyrighted picture with that, you're still infringing. But was the camera created with the intention to appropriate material captured by the lens? Which is why we don't blame the camera for that, we blame the person that used it for that purpose. AI users have an ethical obligation not to steer the AI towards generating infringing material.

Make illegally trained LLMs public domain as punishment in c/technology

[–] ClamDrinker 30 points 16 hours ago (22 children)

Although I'm a firm believer that most AI models should be public domain or open source by default, the premise of "illegally trained LLMs" is flawed. Because there really is no assurance that LLMs currently in use are illegally trained to begin with. These things are still being argued in court, but the AI companies have a pretty good defense in the fact analyzing publicly viewable information is a pretty deep rooted freedom that provides a lot of positives to the world.

The idea of... well, ideas, being copyrightable, should shake the boots of anyone in this discussion. Especially since when the laws on the book around these kinds of things become active topic of change, they rarely shift in the direction of more freedom for the exact people we want to give it to. See: Copyright and Disney.

The underlying technology simply has more than enough good uses that banning it would simply cause it to flourish elsewhere that does not ban it, which means as usual that everyone but the multinational companies lose out. The same would happen with more strict copyright, as only the big companies have the means to build their own models with their own data. The general public is set up for a lose-lose to these companies as it currently stands. By requiring the models to be made available to the public do we ensure that the playing field doesn't tip further into their favor to the point AI technology only exists to benefit them.

If the model is built on the corpus of humanity, then humanity should benefit.

Japan Is So Desperate to Increase Its Birth Rate That Tokyo Is Trying Out a New Idea: Free Daycare in c/[email protected]

[–] ClamDrinker 11 points 1 day ago* (last edited 1 day ago)

Well, countries with higher birthrates have a third option that is essentially negligible in those with lower birthrates, which is not even making it to adulthood. Effectively still less children end up becoming productive members of society. And together with that, due to less available social services, often a goal of having children survive is so they can take care of the parent when they're older.

As soon as infant mortality becomes a non-factor, birthrates decline drastically as well. And since children are no longer largely seen as a "life assurance" for when parents are older, and the society's demands for productive members is higher as well, the focus really does shift to the quality of the life and the two types of reasons to have kids are harder to compare. But even among developed nations you can see differences in fertility rates.

PS. Scandinavia doesn't have the lowest birth rates, they actually have fairly typical birth rates for more developed regions.

Discussing jury nullification is against the ToS of lemmy.world and will get you banned, according to worldnews mod in c/[email protected]

[–] ClamDrinker 4 points 2 weeks ago* (last edited 2 weeks ago)

They're a mod on /c/world and /c/news, just go to the modlist and see it from there. I guess it's not linked directly to avoid brigading.