AI copyright infringement is the unauthorized copying of protected works by artificial intelligence systems, either when those works are ingested to train a model or when a model generates output that reproduces protected expression. Federal law lets a copyright owner recover between $750 and $150,000 per infringed work depending on whether the copying was willful, so the exposure for systems trained on millions of unlicensed books, images, and articles is enormous.1Office of the Law Revision Counsel. 17 US Code 504 – Remedies for Infringement: Damages and Profits Whether that copying is lawful is still being decided case by case, with no appellate court having ruled definitively on large-scale AI training.
How AI Training Can Infringe
Building a large model requires feeding it datasets scraped from the internet. Those datasets routinely contain works whose owners hold the exclusive right to control reproduction and distribution under federal copyright law.2Office of the Law Revision Counsel. 17 US Code 106 – Exclusive Rights in Copyrighted Worksa> During training, the system creates digital copies and processes them to extract statistical relationships between words, pixels, and concepts. Those copies, temporary or not, are reproductions in the legal sense.
The core dispute is whether making those copies without permission violates the reproduction right. Copyright owners say scraping is unauthorized duplication. AI developers say the copies serve a fundamentally different purpose than reading or viewing the original, and that no individual work survives intact inside the finished model. This disagreement drives nearly every major AI copyright lawsuit.
Website terms of service add a second layer of risk for developers. Many sites prohibit automated data collection, and violating those terms can support a breach-of-contract claim separate from any copyright claim. How consistently courts will enforce those provisions against AI companies is still open.
When AI Outputs Infringe
Even if training itself passes legal scrutiny, what a model generates can independently infringe. Courts evaluate output-side infringement using two elements: whether the AI had access to the original work, and whether the output is substantially similar to protected expression in that work.3Ninth Circuit District and Bankruptcy Courts. 17.17 Copying – Access and Substantial Similarity Access is usually easy to establish when the training data is known to include the work. The harder fight is substantial similarity.
Substantial similarity asks whether an ordinary person would recognize the output as taken from the original’s creative expression. The law separates unprotectable elements like general style, genre conventions, and common themes from protectable expression like specific visual compositions, distinctive character descriptions, and unique arrangements. A generated image that replicates the lighting, framing, and color palette of a copyrighted photograph crosses the line even without pixel-perfect copying. An output can also qualify as an unauthorized derivative work if it incorporates recognizable protected elements from a source.4Legal Information Institute. 17 USC 101 – Definitions
One prerequisite catches creators off guard. You generally cannot file a copyright infringement lawsuit in federal court until you have registered the work with the U.S. Copyright Office, or applied and been refused.5Office of the Law Revision Counsel. 17 US Code 411 – Registration and Civil Infringement Actions You still own the copyright without registration. You just cannot enforce it in court without one.
The Fair Use Defense
Fair use is the defense AI companies rely on most, and it is genuinely uncertain whether it will hold up across the range of AI products on the market. Courts weigh four factors, none automatically decisive.6Office of the Law Revision Counsel. 17 US Code 107 – Limitations on Exclusive Rights: Fair Use
Purpose and Character of the Use
The first factor asks whether the use is “transformative,” meaning it serves a different purpose or adds something new rather than substituting for the original. The Supreme Court’s 2021 decision in Google LLC v. Oracle America held that copying code to build a new smartphone platform was transformative because it repurposed the material for a fundamentally different computing environment.7Supreme Court of the United States. Google LLC v. Oracle America, Inc. AI developers lean on that reasoning to argue that ingesting a novel to build a statistical model is a different function than reading or selling it.
The argument has limits. In February 2025, a federal court in Delaware ruled against an AI legal research tool that had trained on copyrighted legal headnotes to build a product competing directly with the copyright holder’s own legal research platform. The court found the use was not transformative because the AI tool served the same purpose as the original product.8United States District Court for the District of Delaware. Thomson Reuters Enterprise Centre GMBH v. Ross Intelligence Inc. The judge did note that generative AI tools might present stronger transformative-use arguments than the non-generative tool in that case. Commercial purpose also weighs against fair use here, though it is not disqualifying by itself.
Nature of the Copyrighted Work
The second factor gives stronger protection to highly creative works like novels, songs, and photographs than to factual compilations. Because AI training data is loaded with creative material, this factor generally favors copyright holders, though courts have treated it as less influential than the others.
Amount Used
AI models typically ingest entire works, which normally weighs against fair use. Courts have recognized that copying the whole work can be justified when the full copy is necessary for the transformative purpose. Whether a model truly needs the complete text of every book it trains on is contested.
Market Effect
This factor carries the most weight. If the AI output substitutes for the original, fair use becomes much harder to establish. When a chatbot summarizes a news article thoroughly enough that readers never visit the publisher’s site, that is direct market harm. The Thomson Reuters court also emphasized the effect on the potential market for AI training data itself, finding that harm to the licensing market weighed against fair use independent of competition in the end product.8United States District Court for the District of Delaware. Thomson Reuters Enterprise Centre GMBH v. Ross Intelligence Inc.
What a Copyright Holder Can Recover
A plaintiff who proves infringement has several remedies, and the arithmetic gets severe fast when a dataset contains millions of works.
- Statutory damages. A copyright owner can elect statutory damages between $750 and $30,000 per work infringed, without proving actual financial harm. Willful infringement raises the ceiling to $150,000 per work. Statutory damages and attorney’s fees are available only when the work was registered before infringement or within three months of publication.1Office of the Law Revision Counsel. 17 US Code 504 – Remedies for Infringement: Damages and Profits5Office of the Law Revision Counsel. 17 US Code 411 – Registration and Civil Infringement Actions
- Actual damages and profits. The owner can instead seek the actual financial harm suffered plus any infringer profits attributable to the infringement.1Office of the Law Revision Counsel. 17 US Code 504 – Remedies for Infringement: Damages and Profits
- Injunctions. Courts can order AI companies to stop using the infringing material, which could mean shutting down or retraining a model.9Office of the Law Revision Counsel. 17 US Code 502 – Remedies for Infringement: Injunctions
- Impoundment and destruction. A court can impound infringing copies and the tools used to produce them during litigation, and a final judgment can order their destruction. Some plaintiffs in AI cases have explicitly asked for the destruction of models trained on their data.10Office of the Law Revision Counsel. 17 US Code 503 – Remedies for Infringement: Impounding and Disposition of Infringing Articles
The window to sue is three years from when your claim accrues.11Office of the Law Revision Counsel. 17 US Code 507 – Limitations on Actions Under the discovery rule applied by most federal circuits, that clock starts when you learn or should have learned about the infringement, not when the copying happened. In 2024, the Supreme Court held that if a claim is timely under the discovery rule, damages can reach back to infringements years before the filing, with no separate cutoff limiting how far back damages extend.12Supreme Court of the United States. Warner Chappell Music, Inc. v. Nealy For creators just now discovering their work was scraped years ago, the clock may not have started until they found out.
Copyright on Work You Made With AI
A related question comes up on the other side. If you used AI to help make something, can you copyright the result? Only to the extent a human authored it. The U.S. Copyright Office requires human authorship for registration.13U.S. Copyright Office. Compendium of US Copyright Office Practices – Section: 302 The Legal Framework In March 2025, the D.C. Circuit affirmed that the Copyright Act requires all eligible work to be authored by a human being.14United States Court of Appeals for the District of Columbia Circuit. Thaler v. Perlmutter Work produced entirely by a machine from a simple prompt falls into the public domain.
AI-assisted work is not automatically unprotectable. If you select, arrange, and edit AI-generated components with real creative judgment, the human-authored portions can qualify. The Copyright Office’s 2023 guidance requires applicants to disclose AI-generated content, describe the human contributions, and exclude more than minimal AI-generated material from the claimed authorship.15U.S. Copyright Office. Copyright Registration Guidance: Works Containing Material Generated by Artificial Intelligence If you already registered a work without disclosing AI use, the Office expects a supplementary registration to correct the record.
The Lawsuits Building the Rules
No appellate court has issued a definitive ruling on whether large-scale AI training qualifies as fair use. A handful of cases are shaping the answer.
Thomson Reuters v. Ross Intelligence
This produced the first federal ruling that AI training on copyrighted material was not fair use. In February 2025, a Delaware judge granted summary judgment to Thomson Reuters after Ross Intelligence used copyrighted legal headnotes to build a competing AI legal research tool.8United States District Court for the District of Delaware. Thomson Reuters Enterprise Centre GMBH v. Ross Intelligence Inc. The court found the use was not transformative because it served the same purpose as the original, and the effect on the potential licensing market for AI training data weighed against fair use.
The New York Times v. OpenAI
The Times sued OpenAI and Microsoft in December 2023, alleging that millions of its articles were scraped to train ChatGPT and that the chatbot now competes with the newspaper by summarizing its journalism.16United States District Court Southern District of New York. The New York Times Company v. Microsoft Corporation, OpenAI, Inc., et al. – Complaint The case puts the market-substitution question front and center. A judge allowed it to proceed in early 2025, and it remains in active discovery.
Getty Images v. Stability AI
Getty accused Stability AI of scraping millions of copyrighted stock photographs to train its image generator, pointing to distorted Getty watermarks appearing in AI-generated images as evidence of direct copying.17Getty Images. Getty Images Statement The U.S. case in Delaware was voluntarily dismissed without prejudice, with Getty indicating plans to refile in the Northern District of California. A UK High Court ruling in late 2025 went in Stability AI’s favor on certain claims, and the litigation continues.
Andersen v. Stability AI
A group of visual artists filed a class action against Stability AI, Midjourney, and DeviantArt over the use of copyrighted artwork to train image generators. Fact discovery is scheduled to close in March 2026, with no dispositive rulings on the copyright claims yet.
Licensing Is Filling the Gap
A parallel licensing market has moved quickly while the courts work. Major AI companies have signed content deals with publishers, news organizations, and media companies. OpenAI has agreements with outlets including Axios, The Guardian, and The Washington Post. Google, Amazon, and Mistral have also entered licensing arrangements with news agencies and publishers. By mid-2025, one AI startup reported that more than 500 publishers had opted into its revenue-sharing program for AI-powered search. The New York Times signed its first AI content licensing deal with Amazon in 2025 while continuing its lawsuit against OpenAI, showing that litigation and licensing can run in parallel.
Practical Steps for Creators
Waiting for the law to settle is not a strategy. A few concrete moves protect your work and preserve your options.
- Register your copyright. You cannot file a federal infringement suit without a registration or pending application. Registering before infringement, or within three months of publication, also unlocks statutory damages and attorney’s fees.5Office of the Law Revision Counsel. 17 US Code 411 – Registration and Civil Infringement Actions
- Block AI crawlers. Add directives to your site’s robots.txt that tell known AI training crawlers not to scrape. OpenAI’s training crawler uses the user-agent string “GPTBot,” and the company publishes a list of its crawlers that respect robots.txt. Robots.txt is not a legal barrier and not every crawler honors it, but it establishes that you did not consent to scraping.18OpenAI. Overview of OpenAI Crawlers
- Check platform terms of service. Some platforms have updated their terms to license user content for AI training. You may need to opt out or move your work.
- Use content credentials. The Coalition for Content Provenance and Authenticity has an open standard for embedding provenance metadata in digital files, creating a verifiable record of who created the content and how it was modified. It doesn’t prevent copying, but it produces an auditable trail.19C2PA. Verifying Media Content Sources
- Monitor AI outputs. Periodically test major AI tools with prompts tied to your work. If a system generates content that closely resembles your protected expression, document it. That evidence supports takedown requests and litigation.
Legislation Is Not Here Yet
Congress has introduced several bills at the intersection of AI and copyright, but none has been enacted as of early 2026. The TRAIN Act would impose transparency requirements on AI companies regarding their use of copyrighted training data.20U.S. Copyright Office. Legislative Developments The NO FAKES Act targets AI-generated replicas of a person’s voice or likeness. The Copyright Labeling and Ethical AI Reporting Act, introduced in February 2026, would add further disclosure obligations. None of these has advanced beyond introduction. The Copyright Act as written, interpreted through the cases moving through the courts, remains the framework you have to work with.