I think Meta is already assuming that there will be no liability for training with copyrighted material. I find it very unlikely that image owners will win the AI training battle.
I'd be extremely surprised if the "Mickey Mouse standing on the moon" example image was a legitimate way to "launder copyright".
The interesting question is just who will be liable for the copyright violation: The party that hosts the AI service? The party that trained it on copyrighted images? The user entering a prompt? The (possibly different) user publishing the resulting image?
Here the problem isn't that the AI was trained on Mickey, but that it generated Mickey. The generated images can still violate copyright if too similar to copyrighted artwork - if published.
I think AI companies are working hard on preventing generated images from being similar to training images unless the user very explicitly asks the result to look like some well known image/character.
You can violate copyright by intentionally drawing Mickey Mouse, the medium of drawing is not relevant (AI can be considered a medium, as much as a digital camera is a medium)
True. This is why I think it's pointless to try to use copyright law to defend yourself against AI companies. Right now, anyway, I don't see any law (or any other mechanism) that provides any protection. If I did, I wouldn't have had to remove all of my websites from the public web.
But is publishing a model which can generate images of Mickey a copyright violation? It's definitely a violation if the model is overfitted to the extent that you can, perhaps lossily, extract the original images.
> But is publishing a model which can generate images of Mickey a copyright violation?
Is selling colored pencils that can draw images of Mickey a copyright violation?
The way I see it, the tool can't ever be at fault for its use, unless its sole use (or something close enough to its sole use) is to infringe in copyright.
Besides, the safeguarding of copyright isn't the single variable we as a society should be solving for. General global productivity is way more valuable than guaranteeing Disney's bottom line.
> The way I see it, the tool can't ever be at fault for its use, unless its sole use (or something close enough to its sole use) is to infringe in copyright.
Even then, you could look at a tape recorder or a photocopier and one of their primary uses is to make a copy of a copyrighted work.
The question isn't "can it be used for" but rather "does it have valid non-infringing use" and "when it does infringe, is it the person who uses the tool or the tool that is at fault?"
> But is publishing a model which can generate images of Mickey a copyright violation?
I don't think that courts have ruled on that specifically (yet), but I seriously doubt that it would be. Taking the image of Mickey and distributing it would certainly be, though.
A photocopier is extremely different because the user is providing the copyrighted material. In the AI case, it is much more like writing a Google search for the copyrighted material. If I ask an artist to draw a cartoon of Mickey Mouse in violation of copyright, the artist is in violation of copyright if they produce said drawing and give it to me. Are we to give special rights to AI that human artists don't enjoy?
When I've heard people talking about using copyright to defend against AI, they've always talked about it in the sense that their works being used to train the AI is where the copyright violation takes place.
That stance is clearly not supported by copyright law.
If, however, we're talking about copyright violations applying to the distribution of works generated by AI, that's an entirely different conversation. It's still not really clear-cut, but there are ways that could be in violation of copyright law.
It isn't the case that AI is being treated differently, though. The issues would be the same if a human were doing all of this stuff.
Tattoo artists also make money off generating infringing content all the time. I thought the issue was not in the generation but in the subsequent usage. Outlawing generation borders on thoughtcrime.
Are tattoo artists breaking the law by creating tattoos of copyrighted material? I think they are. And if an artist becomes really popular for their mickey mouse tattoos, then they will provably be noticed by Disney and there will be consequences.
> The interesting question is just who will be liable for the copyright violation
I don't think this is going to be hard for courts. If you borrow your friends copy of a copyright text, got to kinkos and duplicate it, then distribute the results - you are the one violating copyright, not your friend or kinkos.
The same will hold here I think, mutatis mutandis. This is all completely separable from the training issue.
The person getting sued there would be the user of the model, not meta, as much as I wish that wasn't how it is. If you use photoshop to infringe on copyright, you're at fault, not Adobe.
I don’t agree in this case. Well, maybe I agree on the ultra shitty corporate part. But these are public photos, and if I’d looked at one it could have some influence, probably tiny, on my own drawings. Seems reasonable that the same would be true of my tools.
If they were scanning my private messages, things would be different.
1 - human experience ends up informing human ingenuity. A sketch of Wile E. Coyote comes from someone’s (Chuck Jones?) experience of dogs and seeing coyotes, plus innumerable experience with things that are funny, constraints from experience of certain features that do or don’t work well on animation cels etc. Perhaps a stray tweak in his ears come from a Rembrandt seen as a child or from a glance at a sketch in progress by the person sitting at the next easel in a drawing class long ago.
In todays’s jargon our experiences are all parts of our training set (though today’s massive RNN models are infinitesimal by comparison).
And I think of my tools the same: a ton of inputs stirred together is fine by me.
2 - a difference is that fb’s model is made from public posts: posts offered for anyone to see. In the human case even my private experiences are part of my “training set.”
I don't think any argument in favor of these models that includes reasoning about how humans learn is any good. That's a completely separate process that has very little to do with how these systems work. The issue here is Facebook is creating a commercial system based on data their users have uploaded to their system. If artists had known their work would be used this way, I think they'd rethink using this platform. Facebook's monopoly power over internet content also makes it impractical for you not to have a social media presence if you're trying to make a living as an artist. So you either submit to bullshit like this or damn yourself to obscurity. The fact that it's only training on public content is irrelevant.