This insight first appeared in Bloomberg Law in September 2026 under the title, “DOJ’s OpenAI Copyright Brief Previews AI Licensing’s Future.”
The Bottom Line
- A DOJ statement of interest on AI training and copyright gives transactional lawyers a concrete signal about how the federal government views licensing requirements for training data.
- The OpenAI and Anthropic cases present distinct risks: one centers on the legality of training with lawfully acquired copyrights works. The other involved unauthorized acquisition of third-party content.
- Lawyers advising on AI agreements shouldn’t wait for a final ruling to start conducting due diligence on training data and fair-use risk.

The Department of Justice told a federal judge this month that Congress, not the courts, should decide how artificial intelligence companies pay for the copyrighted material they use to train their models. This position came in a case questioning whether training a large language model on someone else’s journalism is fair use or copyright infringement.
But it’s really an antitrust argument dressed as a copyright defense.
The statement isn’t binding law; it’s a preview of where the administration wants the law to land, filed by the same officials who argue that the decision belongs to somebody else. For anyone drafting or reviewing an AI licensing agreement in the meantime, that preview is worth reading closely — it may be the clearest signal available of how this question gets resolved long before Congress or the courts settle it for good.
Statement of Interest
The filing landed in the US District Court for the Southern District of New York’s multidistrict litigation brought by The New York Times, New York Daily News, and the Center for Investigative Reporting against OpenAI and Microsoft. The DOJ weighed in on the case with a “statement of interest” even though it’s not a party to it. It’s not the type of brief the government files casually; when it does, it brings the full weight of the institution behind it.
This one carries the signatures of associate attorney general Stanley E. Woodward Jr., assistant attorney general Brett Shumate, and senior counsel Michael Weisbuch. They argue that training an AI model on copyrighted text is fair use because the use is transformative — extraordinarily so, in the department’s telling — and that ruling otherwise would hamper the progress of science and the useful arts, the constitutional purpose underlying copyright law itself.
Then the brief makes a more interesting move. Fair use analysis normally asks whether a use harms the market for the original work. The publisher plaintiffs have argued that free training destroys a licensing market they could otherwise build.
The DOJ turns that argument around. Forcing AI companies to license training material, it argued, would let a handful of large media companies set prices that only the biggest AI developers could afford, protecting incumbents on both sides of the transaction and shutting out smaller competitors entirely.
That approach allows the DOJ to support OpenAI’s fair-use position while largely separating the legality of acquiring training data from the question of how that data is used to train a model.
The same brief also tells the court that if publishers want to get paid, the real fix is legislative — a licensing framework Congress could create if it chooses rather than a court ruling that effectively creates one by accident. That’s a defensible position about which branch of government should set this kind of policy.
It’s also an unusual argument to make in a filing whose entire purpose is to persuade a federal judge to resolve the fair use question, on terms the administration prefers, before Congress ever touches it. The DOJ isn’t waiting on Congress to act; it’s working to make sure Congress never has to.
It’s Not Anthropic
The brief’s timing invites a reading that the AI industry is winning across the board. The brief landed five weeks after Anthropic finished paying $1.5 billion to settle a similar fight over its own training data, the largest publicly reported copyright recovery in history. But the comparison mostly demonstrates how different the two cases are.
Anthropic settled after litigation over its downloading of books from pirate sites. The court had already found infringement arising from that piracy, but the settlement did not resolve the broader question of whether AI training itself constitutes infringement when the underlying works were lawfully acquired.
The New York Times doesn’t allege the same kind of acquisition at issue in the Anthropic litigation, where Anthropic downloaded books from known piracy repositories. But acquisition is nevertheless one of the Times plaintiffs’ asserted theories of liability.
As the plaintiffs explained in their summary judgment motion, their claims rest on several distinct alleged bases, including OpenAI’s acquisition of copies of their articles, often from behind paywalls or in violation of industry norms or terms of use.
The case therefore raises both acquisition and downstream-use questions, rather than presenting a simple contrast between Anthropic’s allegedly unauthorized acquisition and OpenAI’s lawful acquisition.
The case is still in discovery and a ruling isn’t imminent, but an appeal seems likely given the stakes involved.
AI Licensing
Lawyers advising on AI vendor contracts, data licensing agreements, and technology mergers and acquisitions shouldn’t just wait for Stein or Congress. The Anthropic settlement and the OpenAI litigation together sketch two vastly different risk profiles for the same underlying activity (training a model on someone else’s content) and diligence needs to be built around telling them apart.
The Anthropic fact pattern is a provenance problem: training data was allegedly acquired through unauthorized copying, including downloading from known piracy repositories. The OpenAI fact pattern presents a more complicated combination of acquisition and use issues.
The Times plaintiffs allege, among other things, that OpenAI acquired copies of their articles without authorization, including material obtained from behind paywalls or in violation of applicable terms or industry norms. The litigation therefore asks not only whether particular uses of copyrighted material were fair, but also whether the manner in which OpenAI obtained that material independently creates liability.
A company can have clean provenance and still face significant fair use risk or messy provenance while avoiding a per-work damages calculation if it can show something closer to authorized access.
That distinction should show up directly in transactional documents. Vendor due diligence questionnaires for any company that trains or fine-tunes models shouldn’t just ask whether training data was licensed. They should ask how it was sourced, whether the vendor can identify the origin of material in its training data, and whether any of that sourcing resembles the bulk downloading at issue in the Anthropic case.
Representations and warranties in AI vendor agreements should be drafted to distinguish provenance risk from use risk rather than folding both into a single generic intellectual property representation. Indemnification provisions negotiated before this litigation concludes should account for the possibility that a favorable outcome for OpenAI narrows fair use risk without eliminating provenance risk, and vice versa if Stein rules the other way.
Companies licensing content to AI developers face a mirror version of the same problem. A publisher or data provider negotiating a licensing deal today is negotiating against a backdrop where the DOJ has told a federal court that requiring licenses at all may be bad policy.
That doesn’t make licensing deals worthless, but it changes their leverage calculus, particularly for smaller content owners who lack the negotiating power of The New York Times. Deal terms that assumed that the courts or Congress would eventually mandate licensing requirements should be revisited with the possibility that neither happens soon.
Deals under negotiation don’t need to wait for a final ruling to account for this uncertainty. Closing conditions and post-closing covenants can be drafted to survive either outcome, tying indemnification caps or purchase price adjustments to defined litigation milestones rather than to a single presumed result.
An AI vendor whose sourcing of training data hasn’t undergone due diligence shouldn’t receive the same representation package as one that has produced a data lineage audit. Buyers in tech company mergers and acquisitions should ask for that audit as a matter of course rather than accepting a boilerplate IP representation written before this litigation existed.
Most-favored-licensing clauses, audit rights over training data provenance, and walk-away rights tied to adverse rulings are all becoming standard requests in sophisticated AI vendor agreements. Counsel who aren’t yet asking for them may be leaving important risks unaddressed.
Making the Rules
The broader point extends beyond this one case. AI policy in the US isn’t being made where the public conversation suggests it should be. It’s being shaped through agency guidance, through insurance companies writing exclusions for AI risks that no statute has defined, and now through a DOJ statement filed in someone else’s lawsuit.
Congress has spent about two years holding hearings on AI. Other parts of the government have spent those same two years deciding what the operative rules will be.

Transactional lawyers should realize that AI risk is being priced and allocated well ahead of any comprehensive statute, in insurance policies, in litigation filed by parties with no stake in the underlying case, and in vendor contracts negotiated without a settled legal baseline to negotiate against.
Waiting for Congress or for a final appellate ruling isn’t a diligence strategy. The companies and counsel positioning themselves well are the ones treating this uncertainty as a known variable to be allocated by contract, not an open question to be resolved later by someone else.
You may also enjoy:
- Another Suit on AI Use of Copyrighted News
- Good Grief: Peanuts Music Owner Sues the Feds and Three Companies for Copyright Infringement
- AI is Now the Witness in Litigation
- AI Product Liability Lawsuit: Grok Deepfake Case Tests Developer Responsibility
and, if you like what you’ve read, please subscribe below or in the right hand column.