The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber.
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
anthropic etal would not have a product to sell without their violation..
kdc had a service that just happened to be popular for pirating...
how are the two even remotely similar?
It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?
The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
This is a too big to fail scenario. If these companies fail, so does the US economy. Normal laws for individuals don't apply, so any comparison to that is pointless.
This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers.
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
It's valid to not take AI companies' side here but people who think publishers are fighing for the little guy's rights are delusional. Tech companies have been exploiting artists for a few years, publishers/record labels/media companies have been doing it for centuries.
I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
This is a settlement that the authors and Anthropic agreed upon.
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
I dont see the relevance. If Anthropic had bought the book at the store, shredded the spine, scanned the pages and trained on that data instead, there wouldnt have been an issue.
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Of course it is. If I write a movie review and sell it to a magazine or whatever, it's derived from the movie, and it's fair use, and I don't need to ask the movie owner for permission first, or give them a cut of my sales. Even if I use some reasonable number of screenshots and video clips, as long as the resulting work is "transformative" i.e. actually a new work, a movie review instead of a copy of the movie.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
>I constantly see people in the AI debate working themselves into wildly copyright maximalist positions
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
IANAL but as an IP creator I have not heard of "derivative products" in the copyright context. There are "derivative works", which are covered by the same copyright as the original. For example, a translation to another language is a derivative work, a novelisation of a movie, a screen adaptation of a book etc. If some author could have proven that any Anthromic model is a derivative work of theirs then they had the copyright on that model and made mad bucks licensing it back to Anthropic.
To sign up for what? The experience of approximately every author on the planet is that they found out that Anthropic did something bad at the same time they were "opted into" the class. The only thing they could do is opt out and litigate on their own against a company with a valuation approaching $1T.
This is a sweet deal for lawyers and for publishers, and nothing else.
Civil justice is primarily about restoring damages, not about punishing wrongdoing (although common law in US it is more punitive than civil law in european countries). Therefore compensations are based on damages, not on profit from wrongdoings.
the irony is that all of this money will go to rent-seeking publishers who won't pass it on to the artists; basically a dispute between the wealthy you're upset with
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
>It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
To keep users paying for content while companies do whatever they want - and if that's not the reason that's certainly an effect.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
> I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
I support anthropics position here, on both learning from and "pirating" books.
The way i see things , the publishers and authors are happy with any policy that makes them more money, and more market control, regardless of what is ethical/just/right.
They would shutdown public libraries , all libraries, if they could.
Aaron Swartz lost his life because he tried to make public knowledge public, and they would be happy to put every information activist to death to protect their monopolies.
IMHO they have no right to stop free access on the internet. The whole copyright system is artificial and monopolistic, and the publishers are complaining yet again, that technology moves information more efficiently than they do, so they want to artificially retard it through goverment action.
The real goverment action that is needed, is to protect private/personal data; not data that is actively traded commercially or publically.
These tech companies are invading personal and private spaces of everyday people, and storing and training with it. Even using it for military targetting and warrantless surveillance.
Anthropic is by no means a good entity, so the way to stick it to them and all tech companies, is to allow their internet scraping, but make it outright criminal to use telemetry or any surveillance techniques they have or will develop.
Also... the ”creators",hollywood,publishers, have no problems scraping themselves, and lift ideas from just about everywhere they can get it.
Almost every hollywood movie is just an assemblage of random memes and topical concerns of everyday ppl, distilled into embelished predictable cheese.
The publishers are the original slop actors. Human Slop.
Aaron swartz lost his life because he committed suicide. Something he had tried multiple times before. If he really only committed suicide because of the legal jeopardy he was in, wouldn't it have made more sense to commit suicide after you're found guilty?
As was pointed out, the settlement is for piracy, not training. They had already ruled that Anthropic's use of copyrighted material for training fell within fair use.
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
> As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead.
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
IANAL, but the parent comment quotes "any profits of the infringer that are attributable to the infringement", which I take to mean it's the profit Anthropic stands to make based on its use of the pirated content that's recoverable.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Because we are mostly discussing a single private person that got caught for maybe 20 songs. I don't want to bring up Aaron but the taste gets saltier the more we see settlements like this.
Well, the USA has jurisdiction over USA companies. If the rest of the world's authors can find a way to obtain jurisdiction over the companies in a way that USA courts won't balk at if asked to enforce, then they're welcome to go ahead.
See, Judge Alsup should have been the person Biden put on the Supreme Court, that or re-nominate Merrick Garland. Instead, he made a silly promise to sate Black Lives Matter, which even when he took office was fast on its way to ignominy, and now Kagan is stuck being the only competent liberal justice on the court. At least Alsup can continue setting the direction of law as it applies to the tech industry.
That depends on the answer to a question that hasn't been answered yet.
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
> That depends on the answer to a question that hasn't been answered yet.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
Good question. Can you ask an LLM to repeat the entire contents of a novel, word-for-word, and read that instead of the original book? I haven't tried it, but I would guess it would not be able to do this.
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
What a fucking joke of a country the US is, allowing this kind of behaviour with such a pathetic "punishment". Barely even qualifies as a tap on the wrist, Anthropic should be getting gutted into non-existence for this shit and the execs should be given the Aaron Swartz treatment.
If you have the time, read the judge's response to the motion:
https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Judge Alsup issued the original order that determined they were liable for piracy but that training LLMs on books was fair use. It's worth reading if you're interested in the topic. https://www.courtlistener.com/docket/69058235/231/bartz-v-an...
Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber.
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
[1] https://www.theverge.com/2017/10/19/16503076/oracle-vs-googl...
But why not jail like Kim Dotcom? And why no Feds jumping on Dario’s window? They are not only pirating, they also resell it!
Yeah. These AI settlements make such a mockery of past copyright enforcement victims that it's straight up offensive.
Police descended upon Kim Dotcom like he was a terrorist or something. They rappelled down helicopters and stormed his home like he was bin Laden.
Then these big techs come along and they make some absurd cost of business settlement.
Kill a man you’re a murderer. Kill a thousand, you’re a conqueror.
Maybe if they could have fined him $1.5B he would have gotten away with it too.
What was Sean Parker sued for again?
Kim built something designed to help everyone pirate stuff. Anthropic pirated specific content.
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
Anthropic pirated that content to help everyone do the same.
Please post the prompts that will reproduce the pirated works verbatim. Or even halfway.
A shitty cam rip of a movie is still punishable as infringement.
That said, it's been done: https://arxiv.org/abs/2601.02671
> In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%).
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
the company is valued at basically 1000x the settlement it is a rounding error for them
copyright infringement was enough to get judgements that ruined entire lives when i was in my late teens and early 20s
now you get to be a founder of a trillion dollar business by extremely large copyright infringement
fuck these ghouls fuck LLMs and fuck the waste of money for this shit
anthropic etal would not have a product to sell without their violation.. kdc had a service that just happened to be popular for pirating... how are the two even remotely similar?
To be clear, the issue is not that the books were used to train Claude, but that they were pirated.
It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
> Fun fact: a pallet of books is called a "gaylord,"
A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
They actually did this.
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
well if they made a PDF copy to process they violated copyright
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
no we should destroy the works of these ghouls and support humans instead of this destructive and useless technology
the people operating frontier labs are bad people they cannot be trusted in any way
the best solution to them would be to send them to monster island (even though it's really a peninsula)
>instead of allowing anyone to train on already scanned books for free
That would be pirating. So your complaint is that they didn't do more piracy?
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
How is this ruled as piracy then? I am confused.
They also pirated the books
They bought, scannned, trained from, and destroyed millions of paper books, which was ruled legal. This lawsuit was for training from LibGen.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
They could have purchased the books instead. It was easier to pirate.
Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?
The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
100% not the case but glad you got paid by them to spread misinformation -- get that paper
Weird to hear a full throated defense of the RIAA here
A summary of what happened is not a full-throated defense of anyone.
naw the person posting that is a paid agent of the RIAA/MPAA
what they are posting is 100% lies
It is wild to see a $1.5 billion resolution in the AI copyright space—especially with around $3,000 per book going directly to affected authors.
This is a too big to fail scenario. If these companies fail, so does the US economy. Normal laws for individuals don't apply, so any comparison to that is pointless.
> If these companies fail, so does the US economy.
I would be really worried about the US economy then.
This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers.
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
https://www.justia.com/entertainment-law/piracy-in-the-enter...
Those are the maximum penalties though
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
It's valid to not take AI companies' side here but people who think publishers are fighing for the little guy's rights are delusional. Tech companies have been exploiting artists for a few years, publishers/record labels/media companies have been doing it for centuries.
If you extrapolate these “fines” to per book piracy, I wonder what the cost would be for something like Anna’s Archive. Trillions?
I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
Edit: grammar
This is a settlement that the authors and Anthropic agreed upon.
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
This case did at least shed light on the fair use argument.
100x the books? Buying a book does not let you redistribute its contents.
If you are selling more than 100 books you are clearly losing out
That ISN’T what this settlement is about? Genuinely please just once read past the headline.
The judge already ruled that training on the books does not constitute reselling their content.
The authors were only owed money for the piracy.
> The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
How many of the authors would license their book for endless creation of derivative works for that amount?
The judge already ruled that it was fair for Anthropic to use books for training if they acquired them legally.
I dont see the relevance. If Anthropic had bought the book at the store, shredded the spine, scanned the pages and trained on that data instead, there wouldnt have been an issue.
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Creating derivative products you charge for surely can't be considered fair use?
>Creating derivative products you charge for surely can't be considered fair use?
All US courts so far have ruled yes.
Of course it is. If I write a movie review and sell it to a magazine or whatever, it's derived from the movie, and it's fair use, and I don't need to ask the movie owner for permission first, or give them a cut of my sales. Even if I use some reasonable number of screenshots and video clips, as long as the resulting work is "transformative" i.e. actually a new work, a movie review instead of a copy of the movie.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
>I constantly see people in the AI debate working themselves into wildly copyright maximalist positions
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
IANAL but as an IP creator I have not heard of "derivative products" in the copyright context. There are "derivative works", which are covered by the same copyright as the original. For example, a translation to another language is a derivative work, a novelisation of a movie, a screen adaptation of a book etc. If some author could have proven that any Anthromic model is a derivative work of theirs then they had the copyright on that model and made mad bucks licensing it back to Anthropic.
YouTubers monetize fair use all the time. Is that significantly different?
Probably few, but irrelevant as the ruling was it was not a derivative work.
> This is a settlement that the authors and Anthropic agreed upon.
The authors or the publishers?
I have a hard time believing they agreed with the millions of authors they pirated.
There were individual authors in the class. They initiated it. Individual authors were allowed to sign up.
If you’re so interested, go read past the headline. Maybe you’ll find that you’re working about what “authors” will agree to.
To sign up for what? The experience of approximately every author on the planet is that they found out that Anthropic did something bad at the same time they were "opted into" the class. The only thing they could do is opt out and litigate on their own against a company with a valuation approaching $1T.
This is a sweet deal for lawyers and for publishers, and nothing else.
Punishable by fine just means it's legal for a cost. If the fine is less than the profit then they'll pay the fine every time.
No the "fine" is 3000 bucks per book.
Thats more than it costs to just shred the spine and scan the book in. Which is probably 15 - 20 bucks a piece.
They will be shredding the book not paying the fine.
Civil justice is primarily about restoring damages, not about punishing wrongdoing (although common law in US it is more punitive than civil law in european countries). Therefore compensations are based on damages, not on profit from wrongdoings.
the irony is that all of this money will go to rent-seeking publishers who won't pass it on to the artists; basically a dispute between the wealthy you're upset with
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
The lawsuit was a mixed blend of individual authors and publishers.
It was started by a group of authors, not publishers.
For the downvotes, my response is to read: https://authorsguild.org/advocacy/artificial-intelligence/wh...
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
>It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
Not true. Individual authors could sign up for the settlement. One of my books was in there under my name.
What's gonna be your payout and are you satisfied with it?
To keep users paying for content while companies do whatever they want - and if that's not the reason that's certainly an effect.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
> I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
Because using pirated material is a civil issue, not a criminal offence?
The laws are for you and me not companies like Anthropic and Meta.
Seems squarely in the "cost of doing business" category
I support anthropics position here, on both learning from and "pirating" books. The way i see things , the publishers and authors are happy with any policy that makes them more money, and more market control, regardless of what is ethical/just/right. They would shutdown public libraries , all libraries, if they could. Aaron Swartz lost his life because he tried to make public knowledge public, and they would be happy to put every information activist to death to protect their monopolies. IMHO they have no right to stop free access on the internet. The whole copyright system is artificial and monopolistic, and the publishers are complaining yet again, that technology moves information more efficiently than they do, so they want to artificially retard it through goverment action. The real goverment action that is needed, is to protect private/personal data; not data that is actively traded commercially or publically. These tech companies are invading personal and private spaces of everyday people, and storing and training with it. Even using it for military targetting and warrantless surveillance. Anthropic is by no means a good entity, so the way to stick it to them and all tech companies, is to allow their internet scraping, but make it outright criminal to use telemetry or any surveillance techniques they have or will develop. Also... the ”creators",hollywood,publishers, have no problems scraping themselves, and lift ideas from just about everywhere they can get it. Almost every hollywood movie is just an assemblage of random memes and topical concerns of everyday ppl, distilled into embelished predictable cheese. The publishers are the original slop actors. Human Slop.
Exactly. Everything is a derivative work, and AI is now making people realise the full extent of that reality.
Aaron swartz lost his life because he committed suicide. Something he had tried multiple times before. If he really only committed suicide because of the legal jeopardy he was in, wouldn't it have made more sense to commit suicide after you're found guilty?
You have to consider that the lawsuit was likely extremely stressful and scary.
Are "open" models exempt from that? Can imagine they used the same datasets.
If they provably shared those datasets then they're just as liable for piracy.
Many of the open weight models are trained on outputs from these models (distillation)
Most open models are developed for profit so they should be equally liable.
that number is missing a zero or two in front of the decimal point
As was pointed out, the settlement is for piracy, not training. They had already ruled that Anthropic's use of copyrighted material for training fell within fair use.
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
(Not that I'm complaining...)
> As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead.
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
[1] 17 U.S.C. § 504(b), https://www.law.cornell.edu/uscode/text/17/504
I would imagine that for over 99% of the books covered in this lawsuit, they're earning less than $3000 per book.
Put another way, their revenues wouldn't drop much if they simply hadn't trained on those 99%.
IANAL, but the parent comment quotes "any profits of the infringer that are attributable to the infringement", which I take to mean it's the profit Anthropic stands to make based on its use of the pirated content that's recoverable.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Thomas-Rasset got 80k per song and Tennenbaum got 22k per song. The law says up to 150k per work. It was a gift.
Sure, but in any other instance of piracy, HN would call awarding $20k per pirated song insane.
Because we are mostly discussing a single private person that got caught for maybe 20 songs. I don't want to bring up Aaron but the taste gets saltier the more we see settlements like this.
So... What about authors from other parts of the world? USA has settled a USA case and they think all is cool for the entire world. So americentric.
Well, the USA has jurisdiction over USA companies. If the rest of the world's authors can find a way to obtain jurisdiction over the companies in a way that USA courts won't balk at if asked to enforce, then they're welcome to go ahead.
> USA has settled a USA case and they think all is cool for the entire world.
Who said that?
Maybe they can give it in the form of expiring Fable credits.
So that the authors can use it to write their next books!
See, Judge Alsup should have been the person Biden put on the Supreme Court, that or re-nominate Merrick Garland. Instead, he made a silly promise to sate Black Lives Matter, which even when he took office was fast on its way to ignominy, and now Kagan is stuck being the only competent liberal justice on the court. At least Alsup can continue setting the direction of law as it applies to the tech industry.
$3k a book is so cheap.
Its probably 100x more than it would have cost to do it legitimately, so seems like reasonable damages to me
That depends on the answer to a question that hasn't been answered yet.
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
> That depends on the answer to a question that hasn't been answered yet.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
Relevant sub-thread here: https://news.ycombinator.com/item?id=48997766
Good question. Can you ask an LLM to repeat the entire contents of a novel, word-for-word, and read that instead of the original book? I haven't tried it, but I would guess it would not be able to do this.
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
I don't think this is plaguarism.
What a fucking joke of a country the US is, allowing this kind of behaviour with such a pathetic "punishment". Barely even qualifies as a tap on the wrist, Anthropic should be getting gutted into non-existence for this shit and the execs should be given the Aaron Swartz treatment.