this post was submitted on 09 Jul 2023

518 points (97.1% liked)

Technology

60950 readers

3889 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each other!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed
Accounts 7 days and younger will have their posts automatically removed.

Approved Bots

founded 2 years ago

MODERATORS

518

2 authors say OpenAI 'ingested' their books to train ChatGPT. Now they're suing, and a 'wave' of similar court cases may follow. (www.businessinsider.com)

submitted 2 years ago by L4s to c/technology

115 comments fedilink hide all child comments

Two authors sued OpenAI, accusing the company of violating copyright law. They say OpenAI used their work to train ChatGPT without their consent.

(page 2) 50 comments

sorted by: hot top controversial new old

[–] [email protected] 4 points 2 years ago* (last edited 2 years ago) (3 children)

The only question I have to content creators of any kind who are worried about AI...do you go after every human who consumed your content when they create anything remotely connected to your work?

I feel like we have a bias towards humans, that unless you're actively trying to steal someone's idea or concepts we ignore the fact that your content is distilled into some neurons in their brain and a part of what they create from that point forward. Would someone with an eidetic memory be forbidden from consuming your work as they could internally reference your material when creating their own?

[–] [email protected] 2 points 2 years ago (14 children)

The problem with AI as it currently stands is that it has no actual comprehension of the prompt, or ability to make leaps of logic, nor does it have the ability to extend and build upon existing work to legitimately transform it, except by using other works already fed into its model. All it can do is blend a bunch of shit together to make something that meets a set of criteria. There's little actual fundamental difference between what ChatGPT does and what a procedurally generated game like most roguelikes do--the only real difference is that ChatGPT uses a prompt while a roguelike uses a RNG seed. In both cases, though, the resulting product is limited solely to the assets available to it, and if I made a roguelike that used assets ripped straight from Mario, Zelda, Mass Effect, Crash Bandicoot, Resident Evil, and Undertale, I'd be slapped with a cease and desist fast enough to make my head spin.

The fact that OpenAI stole content from everybody in order to make its model doesn't make it less infringing.

[–] [email protected] 0 points 2 years ago (1 children)

The fact that OpenAI stole content from everybody in order to make its model doesn’t make it less infringing.

Totally in agreement with you here. They did something wrong and should have to deal with that.

But my question is more about...

The problem with AI as it currently stands is that it has no actual comprehension of the prompt, or ability to make leaps of logic, nor does it have the ability to extend and build upon existing work to legitimately transform it, except by using other works already fed into its model

Is comprehension necessary for breaking copyright infringement? Is it really about a creator being able to be logical or to extend concepts?

I think we have a definition problem with exactly what the issue is. This may be a little too philosophical but what part of you isn't processing your historical experiences and generating derivative works? When I saw "dog" the thing that pops into your head is an amalgamation of your past experiences and visuals of dogs. Is the only difference between you and a computer the fact that you had experiences with non created works while the AI is explicitly fed created content?

AI could be created with a bit of randomness added in to make what it generates "creative" instead of derivative but I'm wondering what level of pure noise needs to be added to be considered created by AI? Can any of us truly create something that isn't in some part derivative?

There’s little actual fundamental difference between what ChatGPT does and what a procedurally generated game like most roguelikes do

Agreed. I think at this point we are in a strange place because most people think ChatGPT is a far bigger leap in technology than it truly is. It's biggest achievement was being able to process synthesized data fast enough to make it feel conversational.

What worries me is that we will set laws and legal precedent based on a fundamental misunderstanding of what the technology does. I fear that had all the sample data been acquired legally people would still have the same argument think their creations exist inside the AI in some full context when it's really just synthesized down to what is necessary to answer the question posed "what's the statically most likely next word of this sentence?"

[–] [email protected] 0 points 2 years ago (1 children)

Is comprehension necessary for breaking copyright infringement? Is it really about a creator being able to be logical or to extend concepts?

I think we have a definition problem with exactly what the issue is. This may be a little too philosophical but what part of you isn’t processing your historical experiences and generating derivative works? When I saw “dog” the thing that pops into your head is an amalgamation of your past experiences and visuals of dogs. Is the only difference between you and a computer the fact that you had experiences with non created works while the AI is explicitly fed created content?

That's part of it, yes, but nowhere near the whole issue.

I think someone else summarized my issue with AI elsewhere in this thread--AI as it currently stands is fundamentally plagiaristic, because it cannot be anything more than the average of its inputs, and cannot be greater than the sum of its inputs. If you ask ChatGPT to summarize the plot of The Matrix and write a brief analysis of the themes and its opinions, ChatGPT doesn't watch the movie, do its own analysis, and give you its own summary; instead, it will pull up the part of the database it was fed into by its learning model that relates to "The Matrix," "movie summaries," "movie analysis," find what parts of its training dataset matches up to the prompt--likely an article written by Roger Ebert, maybe some scholarly articles, maybe some metacritic reviews--and spit out a response that combines those parts together into something that sounds relatively coherent.

Another issue, in my opinion, is that ChatGPT can't take general concepts and extend them further. To go back to the movie summary example, if you asked a regular layperson human to analyze the themes in The Matrix, they would likely focus on the cool gun battles and neat special effects. If you had that same layperson attend a four-year college and receive a bachelor's in media studies, then asked them to do the exact same analysis of The Matrix, their answer would be drastically different, even if their entire degree did not discuss The Matrix even once. This is because that layperson is (or at least should be) capable of taking generalized concepts and applying them to specific scenarios--in other words, a layperson can take the media analysis concepts they learned while earning that four-year degree, and apply them to a specific thing, even if those concepts weren't explicitly applied to that thing. AI, as it currently stands, is incapable of this. As another example, let's say a brand-new computing language came out tomorrow that was entirely unrelated to any currently existing computing languages. AI would be nigh-useless at analyzing and helping produce new code for that language--even if it were dead simple to use and understand--until enough humans published code samples that could be fed into the AI's training model.

[–] [email protected] 1 points 2 years ago

Hmm that is an interesting take.

The movie summary question is interesting. For most people I doubt they have asked ChatGPT for its own personal views on the subject matter. Asking for a movie plot summary doesn't inherrantly require the one giving it to have experienced the movie. If this were the case then pretty much all papers written in a history class would fall under this category. No high schooler today went to war but could write about it because they are synthesizing other's writings about the topic. Granted we know this to be the case and the students are required to cite their sources even when not directly quoting them...would this resolve the first proble?

If we specifically asked ChatGPT "Can you give me your personal critique of the movie The Matrix?" and it returned something along the lines of "Well I cannt view movies and only generate responses based on writings of others who have seen it." would that make the usage more clear? If its required for someone to have the ability to have their own critical analysis, there would be a handful of kids from my high school who would fail at that task too and did so regularly.

I like your college example as that is getting better at a definition, but I think we need to find a very explicit way of describing what is happening. I agree current AI can't do any of this so we are very much talking about future tech.

With the idea of extending matterial, do we have a good enough understanding of how humans do it? I think its interesting when we look at computer neural networks. One of the first ones we build in a programming class is an AI that can read single digit, hand written numbers. What eventually happens is the system generates a crazy huge and unreadable equation to convert bits of an image into a statistically likely answser. When you disect it you'd think, "Oh to see the number 9 the equation must see a round top and a straight part on the right side below it." And that assumption would be wrong. Instead we find its dozens of specific areas of the image that you and I wouldn't necessarily associate with a "9".

But then if we start to think about our own brains, do we actually process reading the way we think we do? Maybe for individual characters. But we know when we read words we focus specifically on the first and last character, the length of the word and any variation of the height of the text. We can literally scramble up the letters in the middle and still read the text.

The reason I bring this up iss that we often focus on how huamsn can transform data using past history but we often fail to explain how this works. When asking ChatGPT a more vague concept it does pull from other's works but one thing it also does is creates a statistical analysis of human speech. It literally figures out what is the most likely next word to be said in the given sentence. The way this calculation occurs is directly related to the matterial provided, the order in which it was provided, the weights programmed into it to make decisions, etc. I'd ask how this is fundamentally different than what humans do.

I'm a big fan of students learning a huge portion of the same literature when in high school. It creates a common dialog we can all use to understand concepts. I, in my 40s, have often referenced a character or event, statement or theme from classic literature and have noticed that only those older than me often get it. In less than a few words I've conveyed a huge amount of information that only occurs when the other side of the conversation gets the reference. I'm wondering if at some point AI is able to do this type of analysis would it be considered transformative?

load more comments (13 replies)

[–] [email protected] 1 points 2 years ago (6 children)

By nature of a human creating something "connected" to another work, then the work is transformative. Copyright law places some value on human creativity modifying a work in a way that transforms it into something new.

Depending on your point of view, it's possible to argue that machine learning lacks the capacity for transformative work. It is all derivative of its source material, and therefore is infringing on that source material's copyright. This is especially true when learning models like ChatGPT reproduce their training material whole-cloth like is mentioned elsewhere in the thread.

load more comments (6 replies)

load more comments (1 replies)

[–] [email protected] 3 points 2 years ago (5 children)

Too be honest, I hope they win. While I my passion is technology, I am not a fan of artificial intelligence at all! Decision-making is best left up to the human being. I can see where AI has its place like in gaming or some other things but to mainstream it and use it to decide who's resume is going to be viewed and/or who will be hired; hell no.

[–] BURN 2 points 2 years ago

I got a degree with a sub focus in AI and I hate where this has gone extremely fast. It’s not exciting anymore, it’s just depressing. I’m trying to get out of tech sooner rather than later and go live off the grid somewhere.

AI will kill society long before it’ll save it

[–] ulu_mulu 1 points 2 years ago (7 children)

I'm not against artificial intelligence, it could be a very valuable tool, but that's nowhere near a valid reason to break laws as OpenAI has done, that's why I too hope authors win.

load more comments (7 replies)

load more comments (3 replies)

[–] [email protected] 2 points 2 years ago

Can’t reply directly to @[email protected] because of that “language” bug, as well. This is an interesting argument. I would imagine that the AI does not have the ability to follow plagiarism rules. Does it even credit sources? I've seen plenty of complaints from students getting in trouble because anti cheating software flags their original work as plagiarism. More importantly I really believe we need to take a firm stance on what is ethical to feed into chat gpt. Right now it's the wild west.

[–] [email protected] 1 points 2 years ago

Good, hope they win.

load more comments