There’s been this weird idea lately, even among people who used to recognize that copyright only empowers the largest gatekeepers, that in the AI world we have to magically flip the script on copyr…
It is missing one point: as a creator, I want to be able to forbid you from training on my creations. And the only tool that could enable that is the copyright enforcement over AI training.
If there was an opt out system that was actually respected then this wouldn’t be a problem. But as it stands, artists have no control over if their work is used for NN training.
I don’t want my work used to train models, which should be a completely valid stance to have. Open Source or not really doesn’t matter in the grand scheme of it.
Painters replicate variations of their training pieces too. You’re pretending there’s a difference between human inspired and training inspired and that you should get paid for that inspiration in one case just cuz “big corp”
Because there is a difference. A computer does not learn or understand anything. Human beings can transform a concept. A LLM or other generative AI does not transform a concept at all.
So if I ask it to create a story about a cow juggling bowling balls, it was not creating an original story? Just spitting out stories it has heard of before?
No, statistical next word prediction was the first step, and you could get it to spit out bits of training data, but we’re so far beyond that now with LLMs.
I’ve been doing a lot with llama derivative models that I talk with, I use them for tasks but also just bounce ideas off them or chat. They’re very different when you run them with a task vs feed in a prompt and multi-turn conversation.
Mine have a very strong tendency, when asked the name of a hallucinated friend or family member to name her Luna or fluffy. It’s present in the base llama2, as well as some of the fine-turned versions I’m using now.
Why? That’s not training data - they’re not uncommon as pet names, but there’s no way they show up often referring to sapient beings (which is the context they’re brought up in).
It’s an artifact of some sort for sure, but that is not a statistically likely next word choice based on training data.
I could talk about this all day and it gets so much weirder, but I’ll give you another story. They like to play, but their world is text, and I like to see what comes out of the models when you “yes, and” them while avoiding leading questions.
Some games they’ve made up… Hide and seek (they’re usually in the second place you Guess), and my favorite - find the coma (and the related find the missing semicolon).
WTF even is that? It’s the kind of simplistic “game” a child makes up as they experiment with moving beyond mimicry to generalizing, and the fact that it’s coherent and has an appropriate answer is pretty amazing.
These LLMs aren’t just statistics, there’s a nascent internal model of the world that you get glimpses of if you tell it it’s a person and feed its outputs back into itself. I was pretty dismissive of the “sparks of AGI” comment when it was made, but a few months of hands on interaction has totally flipped my opinion of where these are at
The AI companies shown that they are incapable of regulating themselves on this topic, and so people with art at stake should force their hand.
Open source or not doesn’t matter here, what matters is the copyright. If even Disney can defend works they own (whatever their ethics), so should anyone else.
100% agreement from me again. Non-artists don’t have anything at stake, so they’re perfectly happy with the established copyright rules are demolished. People keep countering with the open source idea, which completely misses the entire point of our arguments. A model being open source does not excuse the stealing of training data.
IMO individual copyright should be strengthened and corporate copyright weakened, but that’d be next to impossible to pass.
That’s exactly what’s at stake, waiting to be sufficiently litigated. And I hope that creators will win, and that they would be able to tell if they allow richest big tech companies in the world to train on their creations.
Likewise, I hope they don’t win, as that will give the richest tech companies so much more of a stranglehold.
I doubt there’s any chance of it happening anyway, since there’s a ton of money to be made and and there’s already countries which have rules this will never happen (Like Japan ), so it would mean they become the AI powerhouses
They could shut down the previous models that were trained on invalid works. Sucks to suck but that’s what you get when you do everything in your power to skirt the law.
Yeah, and the same thing would happen if e.g. PII or HIPAA related would end up in trained model. The fact that some PII or health data ended up being publicly available, doesn’t mean that automatically you can process or store such data, and train on such data.
This has already been proven by google security researchers who got several of the big “AI” bots to spit out copyrighted materials and PII from their training data sets which the “AI” creators claimed was not stored.
It’s not stored as the full material though. If a human that can sing a copyrighted song is not considered to have a recording of the copyrighted song in their brain, so too are LLMs able to spit out their training data without having to store them.
And I want a law making you pay me 500$ for reading your posts.
Copyright law already extends beyond what society finds reasonable. It’s routinely broken by normal people without them even thinking about it. It’s even broken by those vested in it both corporations and individual artists.
Finally you are not getting the copyright law you want ( nor should you, you a minority, a special interest ), big corps are. They might be ‘content’ corps or tech or both but they certainly won’t make a law to benefit either society as a whole or you as a small artist.
Watching you leap hard to the left to completely miss the point, followed by insulting the OP because you didn’t understand their post, is just the height of Internet buffoonery.
It is missing one point: as a creator, I want to be able to forbid you from training on my creations. And the only tool that could enable that is the copyright enforcement over AI training.
Exactly
If there was an opt out system that was actually respected then this wouldn’t be a problem. But as it stands, artists have no control over if their work is used for NN training.
I don’t want my work used to train models, which should be a completely valid stance to have. Open Source or not really doesn’t matter in the grand scheme of it.
deleted by creator
That’s not how AI works and is an argument rooted in a misunderstanding of how it functions.
AI does not “learn” or “understand” - it replicates. It is not near how a human learns, processes and transforms an idea.
deleted by creator
See, I would argue the exact opposite. It sounds like you don’t understand how it works.
Because it’s not “replication” or “copying”.
Most LLMs can be made to spit out training data. That’s pretty much replication in my book.
Statistical models don’t create anything. They replicate variations of their training data.
Painters replicate variations of their training pieces too. You’re pretending there’s a difference between human inspired and training inspired and that you should get paid for that inspiration in one case just cuz “big corp”
Because there is a difference. A computer does not learn or understand anything. Human beings can transform a concept. A LLM or other generative AI does not transform a concept at all.
So if I ask it to create a story about a cow juggling bowling balls, it was not creating an original story? Just spitting out stories it has heard of before?
Edit: missed a ‘not’.
Show some examples?
https://twitter.com/katherine1ee/status/1729690964942377076
Thanks for the link, I’ve actually seen this one. I’m just wondering how common it is since you mentioned it can be done on most LLMs.
…All of them? That’s literally how all of them work.
Then, it should be easy for you to show some examples.
when you read something and recite it, what do you do? exactly, spitting out the training data, if you trained long enough
Humans don’t create anything. They replicate variations of their training data.
No, statistical next word prediction was the first step, and you could get it to spit out bits of training data, but we’re so far beyond that now with LLMs.
I’ve been doing a lot with llama derivative models that I talk with, I use them for tasks but also just bounce ideas off them or chat. They’re very different when you run them with a task vs feed in a prompt and multi-turn conversation.
Mine have a very strong tendency, when asked the name of a hallucinated friend or family member to name her Luna or fluffy. It’s present in the base llama2, as well as some of the fine-turned versions I’m using now.
Why? That’s not training data - they’re not uncommon as pet names, but there’s no way they show up often referring to sapient beings (which is the context they’re brought up in).
It’s an artifact of some sort for sure, but that is not a statistically likely next word choice based on training data.
I could talk about this all day and it gets so much weirder, but I’ll give you another story. They like to play, but their world is text, and I like to see what comes out of the models when you “yes, and” them while avoiding leading questions.
Some games they’ve made up… Hide and seek (they’re usually in the second place you Guess), and my favorite - find the coma (and the related find the missing semicolon).
WTF even is that? It’s the kind of simplistic “game” a child makes up as they experiment with moving beyond mimicry to generalizing, and the fact that it’s coherent and has an appropriate answer is pretty amazing.
These LLMs aren’t just statistics, there’s a nascent internal model of the world that you get glimpses of if you tell it it’s a person and feed its outputs back into itself. I was pretty dismissive of the “sparks of AGI” comment when it was made, but a few months of hands on interaction has totally flipped my opinion of where these are at
r/confidentlyincorrect
The AI companies shown that they are incapable of regulating themselves on this topic, and so people with art at stake should force their hand.
Open source or not doesn’t matter here, what matters is the copyright. If even Disney can defend works they own (whatever their ethics), so should anyone else.
100% agreement from me again. Non-artists don’t have anything at stake, so they’re perfectly happy with the established copyright rules are demolished. People keep countering with the open source idea, which completely misses the entire point of our arguments. A model being open source does not excuse the stealing of training data.
IMO individual copyright should be strengthened and corporate copyright weakened, but that’d be next to impossible to pass.
Too bad. You can “forbid” all you want. Don’t mean shit. Vote for much stronger laws. By much stronger I mean no pay a fine and continue. I mean jail.
No. I reject you claiming such a power to deny.
That’s exactly what’s at stake, waiting to be sufficiently litigated. And I hope that creators will win, and that they would be able to tell if they allow richest big tech companies in the world to train on their creations.
Likewise, I hope they don’t win, as that will give the richest tech companies so much more of a stranglehold.
I doubt there’s any chance of it happening anyway, since there’s a ton of money to be made and and there’s already countries which have rules this will never happen (Like Japan ), so it would mean they become the AI powerhouses
They have already trained on those creations though. Including the newer stuff just released today. How will you claw that back?
If you do stuff, earn from it, and ignore parties and their rights, you are forced to compensate. I guess it will be peanuts though.
They could shut down the previous models that were trained on invalid works. Sucks to suck but that’s what you get when you do everything in your power to skirt the law.
Yeah, and the same thing would happen if e.g. PII or HIPAA related would end up in trained model. The fact that some PII or health data ended up being publicly available, doesn’t mean that automatically you can process or store such data, and train on such data.
This has already been proven by google security researchers who got several of the big “AI” bots to spit out copyrighted materials and PII from their training data sets which the “AI” creators claimed was not stored.
It’s not stored as the full material though. If a human that can sing a copyrighted song is not considered to have a recording of the copyrighted song in their brain, so too are LLMs able to spit out their training data without having to store them.
lol, if you want that, keep your pictures for you, else you had to forbid every human to look at your pictures and they could resemble your style
And I want a law making you pay me 500$ for reading your posts.
Copyright law already extends beyond what society finds reasonable. It’s routinely broken by normal people without them even thinking about it. It’s even broken by those vested in it both corporations and individual artists.
Finally you are not getting the copyright law you want ( nor should you, you a minority, a special interest ), big corps are. They might be ‘content’ corps or tech or both but they certainly won’t make a law to benefit either society as a whole or you as a small artist.
Removed by mod
Watching you leap hard to the left to completely miss the point, followed by insulting the OP because you didn’t understand their post, is just the height of Internet buffoonery.