The Justice Division has an issue with Choose Vince Chhabria’s opinion in Kadrey v. Meta. Really, it has two issues.
First, DOJ likes Choose Chhabria’s conclusion that coaching a big language mannequin on copyrighted books might be extremely transformative. The Division cites Kadrey approvingly for precisely that proposition. (I don’t purchase it, however there it’s.)
Second, DOJ very a lot doesn’t like what Choose Chhabria stated subsequent — which is probably the most cheap, and subsequently my favourite, a part of Choose Chhabria’s opinion.
In its (weird) Assertion of Curiosity within the New York Instances v. OpenAI litigation, DOJ characterizes Choose Chhabria’s dialogue of market dilution as “[c]ontrary dicta” that “misapplies copyright rules to LLM coaching.” Assertion of Curiosity of america at 14, New York Instances Co. v. Microsoft Corp., No. 1:23-cv-11195 (S.D.N.Y. 2025) (DOJ Assertion). Based on DOJ, Choose Chhabria improperly “collapsed LLM coaching and LLM outputs right into a single steady use.” Id. at 15. Coaching, DOJ insists, is “its personal use.” Id. Coaching itself substitutes for nothing, and aggressive outputs typically shouldn’t depend as cognizable market hurt until they reproduce considerably related protected expression.
That description obscures what Choose Chhabria really stated.
What Did Choose Chhabria Really Say?
Choose Chhabria agreed that Meta’s use of copyrighted books to coach Llama was “extremely transformative.” Kadrey v. Meta Platforms, Inc., No. 23-cv-03417-VC, slip op. at 27 (N.D. Cal. 2025). The books weren’t being copied merely to republish them. They had been getting used to develop a know-how able to performing capabilities totally different from these carried out by the books themselves.
However Choose Chhabria didn’t cease the fair-use evaluation there. His concern was the distinction in sort and scale between earlier transformative-use circumstances and generative AI. Id. at 32:
This case is totally different. This isn’t a case the place an authentic work is being in comparison with one secondary work. Neither is this case just like the earlier honest use circumstances involving creation of a digital software. In these circumstances, like Google Books and Excellent 10, the software may at most be used to entry half or the entire authentic works. This case, in contrast to any of these circumstances, entails a know-how that may generate actually thousands and thousands of secondary works, with a miniscule fraction of the time and creativity used to create the unique works it was educated on. No different use — whether or not it’s the creation of a single secondary work or the creation of different digital instruments — has something close to the potential to flood the market with competing works the way in which that LLM coaching does.
And so the idea of market dilution turns into extremely related, and that is a vital distinction.
Clearly so.
Choose Chhabria was not discussing similarity of “fashion” or that each output from an LLM infringes the works used to coach it. His pointed concern was market impact — the fourth issue Congress expressly instructed courts to contemplate.
A know-how could also be terribly transformative and terribly disruptive of the marketplace for the works that made the transformation attainable. These propositions aren’t contradictory. That’s the reason §107 has 4 elements as a substitute of 1.
The Kadrey plaintiffs however misplaced. Why? Id. at 33–34:
On this case, as a result of Meta’s use of the works of those 13 authors is very transformative, the plaintiffs wanted to win decisively on the fourth issue to win on honest use. See, e.g., Excellent 10, 508 F.3d at 1168 (honest use the place secondary use was “important[ly] transformative” and fourth issue “favor[ed] neither social gathering”). And to stave off abstract judgment, they wanted to create a real difficulty of fabric reality as to that issue. As a result of the problem of market dilution is so necessary on this context, had the plaintiffs introduced any proof {that a} jury may use to search out of their favor on the problem, issue 4 would have wanted to go to a jury. Or maybe the plaintiffs may even have made a powerful sufficient exhibiting to win on the honest use difficulty at abstract judgment. However the plaintiffs introduced no significant proof on market dilution in any respect. Absent such proof and in mild of Meta’s proof, the fourth issue can solely favor Meta. Subsequently, on this document, Meta is entitled to abstract judgment on its honest use protection to the declare that copying these plaintiffs’ books to be used as LLM coaching information was infringement.
That doesn’t make Choose Chhabria’s reasoning disappear regardless of how a lot DOJ desires a pony.
The Dangerous Hemingway Contest at Fundamental Justice
DOJ tries to make Choose Chhabria’s reasoning sound absurd by invoking Joan Didion. As DOJ tells it, Didion reportedly typed out Hemingway tales as an adolescent to find out how his sentences labored. If Choose Chhabria had been right, DOJ suggests, maybe Didion ought to have needed to compensate Hemingway each time she subsequently revealed literature that competed together with his.
[W]hen she was an adolescent, Joan Didion “would kind out” Ernest Hemingway’s “tales to find out how the sentences labored,” and because of this she thought of him the best affect on her writing… By the Kadrey courtroom’s logic, Didion ought to have incurred legal responsibility to Hemingway each time she revealed a chunk, as a result of the method by which she educated herself and the method by which she produced works was all one use, and her works competed with these of different authors out there for literature. However “to make anybody pay particularly for the usage of a guide . . . every time they later draw upon it when writing new issues in new methods could be unthinkable.”
(Citing Bartz.) DOJ Assertion at 16–17.
It’s a memorable analogy, however not for the meant cause. In truth, it’s not a lot of an analogy in any respect.
Didion was a human writer finding out one other human writer. She didn’t ingest Hemingway’s collected works right into a business machine able to producing thousands and thousands of competing books on demand.
Extra basically, the analogy assumes away the authorized query. Didion didn’t have to make industrial-scale machine copies of copyrighted works to own a mind able to studying from them. AI builders do make copies. That’s the reason §107 is concerned within the first place. That’s sort of the entire level. In truth, it’s the complete level.
Calling each actions “coaching” no extra makes them legally equal than calling each Didion and Llama “writers” makes them technologically equal. Or calling an orange a pomegranate.
And if the proposition is merely that human beings have all the time discovered to mimic writers, DOJ may simply as simply have cited the previous Harry’s Bar Dangerous Hemingway Contest. No person significantly thought the contestants owed Hemingway royalties as a result of they may produce dangerous Hemingway. Joan Didion was not an AI mannequin. Harry’s Bar was not a knowledge heart.
The copyright query begins the place the joke ends: how the machine acquired the potential, at what scale, and what occurs to the market when that functionality is commercialized. And had the Kadrey plaintiffs placed on a strong market hurt case, we may be having a wholly totally different dialog (or if Meta hadn’t constructed a enterprise buying and selling on kids’s vulnerabilities for that matter).
The place Is DOJ’s Limiting Precept?
This brings us to the bigger downside with DOJ’s place. The issue was there. It was a big downside and DOJ knew it was there. They seemed for a limiting precept. There was no limiting precept. The solar was sizzling and the briefs had been very lengthy.
Or, to place it much less like an entry within the Dangerous Hemingway Contest: there doesn’t look like a limiting precept. Presumably as a result of there isn’t one.
DOJ says coaching is awfully transformative as a result of copyrighted works are copied to assemble an LLM able to doing one thing new. However when DOJ reaches issue 4, the ensuing functionality immediately disappears from the evaluation. Based on DOJ, coaching (or inputs) is one use. Outputs are one other. Coaching itself substitutes for nothing. DOJ would have you ever imagine that the coaching copies aren’t uncovered to shoppers (chatbots however). And outputs that compete with authors supposedly don’t depend until they reproduce considerably related protected expression.
Observe that reasoning to its conclusion. What unauthorized copying for AI coaching may ever fail honest use?
The developer might copy a complete work. That presents little problem as a result of shoppers supposedly by no means see the coaching copy. The AI developer might copy thousands and thousands of works for a business product. Commerciality carries comparatively little weight as a result of the use is very transformative. The developer might use these copies to assemble a machine able to producing competing expression. That doesn’t depend as a result of creating the potential is transformative. The ensuing machine might then generate thousands and thousands or billions of works competing in exactly the markets occupied by the individuals whose works had been copied. That doesn’t depend as a result of outputs are a separate “use.” Even when it’s additionally referred to as Suno.
What stays? Primarily, regurgitation. However regurgitation is an output infringement downside. DOJ has constructed a rule below which the antecedent copying is successfully insulated from the very financial penalties for which the copying was undertaken.
That appears remarkably like an AI-training protected harbor. And I believe that’s precisely what DOJ is attempting to perform as a substitute of prosecuting what is clearly… clearly… probably the most intensive act of prison copyright infringement of all time.
Inputs and Outputs Are Not Strangers
There may be additionally one thing basically illogical about DOJ’s insistence on separating inputs from outputs. One implies the opposite. Clearly so. The extra you concentrate on it, the extra apparent the fallacy turns into. In truth, “clearly” is the go-to adverb in AI. You already know, clearly progressively at first, then immediately. Or one thing like that.
The copyrighted works are copied as a result of the developer desires the mannequin to accumulate capabilities that subsequently manifest themselves via outputs. Coaching will not be an finish in itself. No person spends billions of {dollars} constructing an LLM in order that it may sit quietly considering its weights. The output functionality is the product. It’s the entire level. Clearly.
Certainly, DOJ depends on precisely this causal relationship when discussing transformation. The copied works are remodeled into numerical representations; the mannequin learns linguistic relationships from them; the ensuing system can then reply to customers (assuming it’s been educated on thousands and thousands of books in lots of of languages, which it must be to ensure that the platform to launch a worldwide product).
That relationship helps DOJ below issue one as a result of it demonstrates that the aim of the copying differs from the aim of the books. Then issue 4 arrives. All of a sudden the output functionality turns into anyone else’s downside.
The asymmetry is putting:
- Enter → output functionality counts when proving transformation.
- Enter → output competitors doesn’t depend when proving market hurt.
That isn’t a use-by-use evaluation. It’s a factor-by-factor manipulation of the boundary of the system’s limiting precept.
Google Books Did Not Work That Means
The artificiality turns into clearer when put next with Authors Guild v. Google.
Google copied thousands and thousands of books. Choose Pierre Leval’s opinion on attraction discovered the copying transformative as a result of the ensuing database allowed customers to seek for data about books and find responsive passages. Choose Leval described Google’s copying as transformative as a result of:
Google’s making of a digital copy to supply a search perform is a transformative use, which augments public information by making out there details about Plaintiffs’ books with out offering the general public with a considerable substitute for matter protected by the Plaintiffs’ copyright pursuits within the authentic works or derivatives of them.
Authors Guild v. Google, Inc., 804 F.3d 202, 207 (second Cir. 2015)(emphasis mine).
That final qualification wasn’t an afterthought. It was constructed into Choose Leval’s description of why the use was honest.
And Choose Leval expressly linked the inputs to what the ensuing machine may do. The “digital corpus created by the scanning of those thousands and thousands of books,” he defined, “permits the Google Books search engine.” Id. at 209. The corpus additionally enabled “new types of analysis, referred to as ‘textual content mining’ and ‘information mining,’” permitting researchers to look at “phrase frequencies, syntactic patterns, and thematic markers” throughout tens of thousands and thousands of books. Id.
In different phrases, the outputs defined why the inputs had been copied. And people outputs mattered to issue 4, too.
Though I by no means believed it, Choose Leval particularly examined what Google really gave customers. His conclusion was not merely that Google’s ingestion was transformative and subsequently the evaluation was over. The courtroom held that Google’s copying “doesn’t provide the general public a significant substitute for matter protected by the plaintiffs’ copyrights.” Id. at 222–23. The opinion repeatedly returns to that limitation.
Snippet View was subsequently intentionally constrained. Google completely blacklisted parts of each guide, restricted the snippets returned for a search, prevented customers from increasing disclosure via repeated searches, and disabled snippet view totally for works — comparable to cookbooks, dictionaries and books of quick poems — the place even a small excerpt may fulfill the person’s want for the unique. Id. at 222–24.
These restrictions mattered as a result of what got here out of the system mattered. Google Books subsequently didn’t stand for the proposition that when ingestion is transformative, courts avert their eyes from what the ensuing machine does. Fairly the reverse. The machine’s performance helped set up transformation and its limitations helped set up the absence of substitution.
In any other case we’d say Google copied thousands and thousands of books merely to assemble the digital Library of Alexandria, whereas declining to acknowledge what anyone meant to do with the corpus afterward. Corpus machine translation anybody? Oh, sorry.
Choose Leval Already Equipped the Limiting Precept
There may be further irony right here. DOJ’s argument is troublesome to reconcile not solely with Choose Leval’s seminal and influential regulation evaluation article, Towards a Truthful Use Commonplace, from which trendy transformative-use doctrine largely emerged, however with Choose Leval’s personal software of that doctrine within the Second Circuit Google Books opinion.
Choose Leval started the Court docket’s Google Books opinion with an announcement of info that I believe facially distinguishes Google Books from no matter world the DOJ occupies (emphasis mine). Id. at 207–208:
By way of its Library Challenge and its Google Books challenge, appearing with out permission of rights holders, Google has made digital copies of tens of thousands and thousands of books, together with Plaintiffs’, that had been submitted to it for that objective by main libraries. Google has scanned the digital copies and established a publicly out there search perform. An Web person can use this perform to go looking with out cost to find out whether or not the guide incorporates a specified phrase or time period and in addition see “snippets” of textual content containing the searched-for phrases. As well as, Google has allowed the collaborating libraries to obtain and retain digital copies of the books they submit, below agreements which commit the libraries to not use their digital copies in violation of the copyright legal guidelines. These actions of Google are alleged to represent infringement of Plaintiffs’ copyrights.
In different phrases, neither the Library Challenge nor the Google Books challenge bore any resemblance by any means to synthetic intelligence fashions or OpenAI.
Furthermore, NYT v. OpenAI bears no resemblance to Authors Guild v. Google. In Google Books, the Second Circuit discovered honest use as a result of it discovered that Google’s system was structurally incapable of producing new expressive content material — it listed copyrighted works and returned snippets as pointers again to the originals, functioning as a complementary discovery software that drove customers towards buying or borrowing books and posed no market hurt. (Did it additionally kind the muse for Google’s LLM?). OpenAI’s fashions are categorically totally different: they ingest copyrighted journalism throughout coaching, internalize its substance into mannequin weights, after which generate outputs that serve the identical informational perform the unique reporting served — delivering the Instances’s analytical and investigative worth to customers who would in any other case have learn the Instances, with out sending them to the Instances or compensating it.
Whether or not any specific output is a verbatim copy is, to me, inappropriate; the mannequin has absorbed the worth of the work, and the aggressive hurt lies in practical substitution, not character-for-character copy. Worse nonetheless, OpenAI’s system is generative by design, able to flooding the market with an primarily limitless quantity of competing content material on the identical matters, in an analogous register, at near-zero marginal value — industrialized substitution at a scale that devalues the originals exactly as a result of it replicates their informational perform throughout thousands and thousands of queries per day, eroding the financial basis of the very journalism it was educated on.
I believe we are able to generalize fairly safely that the identical will probably be true of works in any copyright class, or personhood itself.
As Choose Leval wrote:
…the Supreme Court docket has made clear that among the statute’s 4 listed elements are extra important than others. The Court docket noticed in Harper & Row Publishers, Inc. v. Nation Enterprises that the fourth issue, which assesses the hurt the secondary use could cause to the marketplace for, or the worth of, the copyright for the unique, “is undoubtedly the one most necessary component of honest use.” 471 U.S. 539, 566 (1985) (citing Melville B. Nimmer, 3 Nimmer on Copyright § 13.05[A], at 13–76 (1984)). That is in keeping with the truth that the copyright is a business proper, meant to guard the flexibility of authors to revenue from the unique proper to merchandise their very own work.
Authors Guild, 804 F.3d at 212.
It ought to come as a shock to nobody that’s totally constant together with his earlier formulation in Towards a Truthful Use Commonplace. Choose Leval didn’t say that discovering a transformative objective ends the evaluation. His formulation was significantly extra cautious. A good use should additional copyright’s goal of stimulating productive thought and public instruction “with out excessively diminishing the incentives for creativity.” Leval, Towards a Truthful Use Commonplace, 103 Harv. L. Rev. 1105, 1110 (1990).
And when Choose Leval turned particularly to the fourth issue, the connection was specific: “A secondary use that interferes excessively with an writer’s incentives subverts the goals of copyright.” Id. at 1124.
That’s the limiting precept.
Transformation issues as a result of copyright exists for the general public profit, not merely to guard incumbent markets. However transformation can’t be permitted to devour the financial situations needed for continued creation.
And Google Books illustrates the stability significantly effectively. Choose Leval didn’t merely conclude that Google had invented one thing helpful. He concluded that the helpful factor Google invented didn’t present the general public with a “substantial substitute” for the protected expression it had copied. Id. at 207.
Generative AI presents the query from the other way. The copied corpus will not be merely being searched. It’s getting used to assemble a system whose business objective is to generate new expression. That doesn’t mechanically make coaching unfair. However neither can honest use coherently declare the ensuing expressive functionality related when asking whether or not coaching is transformative and irrelevant when asking what the coaching does to the market.
The DOJ’s assertion is past dangerous Hemingway. It’s only a elementary misreading of the info, the regulation, and, I might argue, each precept that Choose Leval enunciated in his intensive and authoritative honest use discussions.
The Firewall within the Causal Chain
Choose Chhabria didn’t “collapse” two unrelated makes use of in Kadrey, because the DOJ accuses. He adopted the causal chain — and rightly so. Copyrighted works are copied into coaching datasets. The mannequin operationalizes these copies, extracting expressive capabilities throughout a spread of functionalities. These capabilities produce outputs. And people outputs enter markets at a scale by no means earlier than seen, at zero or near-zero marginal value per copy. This isn’t analytical confusion — it’s analytical precision.
The coaching and the output aren’t unbiased occasions to be assessed in isolation; they’re sequential steps in a single pipeline that begins with mass ingestion of copyrighted materials and ends with the business distribution of competing works derived from it. To sever the chain — to guage the copying with out contemplating what it permits, or to guage the output with out tracing it again to its supply — is to disregard the financial actuality of how these programs really perform.
Choose Chhabria acknowledged that the honest use inquiry should account for your complete arc from ingestion to market impression, as a result of that’s the place the hurt materializes. The DOJ’s insistence on disaggregating coaching from output will not be doctrinal rigor; it’s a framework designed to make the copying disappear from the evaluation exactly in the intervening time it issues most.
Issue one asks why the works had been copied.
Issue 4 asks what occurs if copying for that objective turns into widespread.
These questions look at totally different facets of the identical challenged conduct. Treating the financial consequence of coaching as categorically irrelevant as a result of it manifests itself via the know-how that coaching created doesn’t faithfully apply §107. It prevents issue 4 from doing the job Congress assigned it. And that’s the reason DOJ’s place has no significant stopping level.
If wholesale copying is transformative due to the capabilities it creates, whereas the market penalties of these capabilities can’t be attributed to the copying as a result of outputs are a separate use, nearly each non-regurgitative coaching use begins to look honest by definition. Choose Chhabria refused to create that rule. He as a substitute utilized a a lot older one: courts ought to ask what occurs to authors and their markets if everyone seems to be allowed to do what the defendant has completed.
That’s not some novel AI exception to honest use. It’s the conventional market inquiry utilized to a know-how succesful, as Choose Chhabria put it, of producing “actually thousands and thousands” of competing works. Kadrey, slip op. at 32.
DOJ’s reply is to interrupt the causal chain at precisely the purpose the place persevering with to comply with it turns into inconvenient.
That’s not a limiting precept.
It’s a firewall erected in the course of the causal chain.
The Report Is the Lesson
Choose Chhabria’s evaluation in Kadrey was doctrinally sound. It was in keeping with Google Books, in keeping with Choose Leval’s foundational work on transformative use, and in keeping with the statutory construction Congress enacted. He didn’t collapse two unrelated makes use of. He didn’t invent a novel AI exception. He utilized the standard four-factor inquiry to a know-how that occurs to function at unprecedented scale — and he concluded, accurately, that market dilution is very related when the challenged copying permits a machine to generate thousands and thousands of competing works.
The Kadrey plaintiffs misplaced not as a result of the framework failed them, however as a result of the document did. They introduced no significant proof of market hurt. Choose Chhabria stated so plainly, and he was equally plain that the result might need been considerably totally different had the evidentiary basis been there.
That’s the forward-looking lesson of Kadrey — not that honest use must be reinvented for synthetic intelligence, however that in copyright circumstances market hurt should be confirmed with the identical evidentiary rigor courts have all the time required. The framework is sound. The causal chain is actual. What stays is for plaintiffs to construct the document that follows it to its conclusion.
Within the phrases of the person himself, “Essentially the most important reward for a superb author is a built-in, shockproof, shit detector.” — Ernest Hemingway, The Paris Evaluation Interview (1958).




