A new study from MIT computer scientists published in Nature Communications on Tuesday delivers a finding that could reshape the debate over AI copyright: as diffusion models are trained on more data, their outputs become increasingly impossible to trace back to any specific piece of training material. The researchers call the phenomenon "attribution decay." For readers tracking the fast-moving world of generative AI, the paper is one of the more consequential research results covered this week by AI Buzz Wire.

The paper, titled "Outputs of Generative Diffusion Models are Often Unattributable," was authored by Zheng Dai and David K. Gifford, both affiliated with MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL). Their central question was deceptively simple: when a generative model produces an image, can you locate a specific item in the training data that is responsible for that output?

What the Researchers Found

The answer, increasingly, is no. The researchers show that attribution — defined as the task of locating part of the training data that can be held responsible for a generated sample — "can become impossible if a model is trained on a sufficiently large corpus of data."

"We find that the more data a model is trained on, the less attributable its generated samples become, a phenomenon we henceforth refer to as attribution decay," the authors write.

To test this, the team developed an "ablation" methodology that allows specific training examples to be removed from an already-trained model without costly retraining. They then ran large-scale what-if experiments: what would the model generate if a particular image, or an entire creator's body of work, had never been in the training set?

In some of the study's most striking results, the researchers found that for models trained on sufficiently large datasets, they could remove the Mona Lisa — or even all of Leonardo da Vinci's work — and the model could still reproduce the image or the artistic style. No single training item was causally responsible for the output.

Why It Matters for Copyright Litigation

The findings land in the middle of an active legal battlefield. Diffusion models such as Midjourney and Stable Diffusion have attracted lawsuits from artists who argue that copies of their work in training data allow AI systems to reproduce their output and style.

In one ongoing copyright case from 2023, Andersen et al. v. Stability AI Ltd., plaintiffs have been trying to force Midjourney to disclose the datasets used to train its models. They allege Midjourney built its training sets "by scraping images associated with specific artists' names for the express purpose of enabling its model to mimic those artists' expressive content."

The MIT study suggests that establishing such causal links may be technically impossible at scale. As The Register, which first reported the study's details, noted, the research looks likely to make AI regulation more difficult rather than easier — the opposite of what the authors hoped when they began the work.

The Fair Use Question

In MIT's press release, Gifford argues the findings cut both ways. If generated outputs have nothing to do with any individual piece of training data, models may be genuinely creative rather than mere copying machines.

"That raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn't attributable to anything on the internet," Gifford said.

The implications extend beyond image generators. Diffusion models have become the dominant architecture for generating audiovisual media and are increasingly used in scientific applications, including protein structure modeling and therapeutic discovery. Attribution tools have also been proposed for machine unlearning, data poisoning detection, model interpretability, fairness audits, and privacy compliance.

A Warning for Attribution Startups

The study also carries a caution for the growing industry of AI provenance and content-verification tools. The researchers found that similarity-based attribution — matching generated outputs to visually similar training images — produces false attributions in large training data regimes.

In other words, a tool that claims "this AI image came from that artist's portfolio" may simply be wrong when models are trained on billions of images. For platforms, courts, and regulators counting on technical attribution to settle provenance disputes, the MIT results suggest such certainty may be out of reach for frontier-scale models.

What Comes Next

The paper is expected to be cited in ongoing litigation and policy debates, including discussions around the EU AI Act's transparency requirements and proposals for artist compensation systems. If attribution decay holds for the largest production models, compensation schemes based on tracing outputs to specific works may need to be redesigned around collective licensing rather than item-level tracking.

The study also raises a subtle point for AI developers: unattributability may actually strengthen fair use arguments for training on copyrighted data, while simultaneously undermining artists' ability to prove specific harm from individual outputs. Both sides of the copyright fight will find something to use in the data.

Stay Ahead of AI

For more breaking AI research and analysis, visit AI Buzz Wire for daily coverage of the AI industry.

Read more AI news →