18 August 2026
MIT researchers find AI images often untraceable to any single training source
- MIT CSAIL researchers discovered attribution decay, a phenomenon where large AI image generators become increasingly disconnected from individual training images as dataset size grows.
- The team built a diffusion ensemble, a new architecture made of smaller components instead of one large model, allowing them to test what would happen if specific training images were removed without retraining from scratch.
- Testing on datasets from 256 to 160,000 images showed a consistent pattern: the larger the training set, the less any single image affected the final output, following an inverse power law.
- If removing a training image produces no change in output, the researchers argue that image cannot be said to have contributed to that output, complicating legal questions about whether AI-generated images are derivative works.
- The finding applies to diffusion models used in image generation and scientific work like protein modeling, though it remains unclear whether the same decay occurs in large language models at the center of copyright lawsuits.