
Output attribution cannot replace the training ledger
A diffusion study separates an output's causal attribution from its training inputs. Keep the input record and output investigation as separate checks; neither certifies the other.
Removing a training image need not make a generated image noticeably different. That does not mean the training image was never used.
Zheng Dai and David Gifford's August 18 Nature Communications study makes that distinction concrete. Across 24 specially constructed diffusion ensembles and seven image datasets, the researchers found that individual outputs became less attributable as training sets grew. Their test asks how much an output changes when a unit of training data is removed, keeping generation conditions fixed. A unit can be an image or a group associated with a person or artist.
The scope matters: these are bespoke ablatable ensembles, not an audit of a commercial model. The finding does not establish that language models behave identically, or that training data stops mattering. It also does not rule out useful attribution to collections of data.
My rule for a model review is simpler than the research: never let an output-attribution report answer an input-inventory question.
Two questions, two receipts
An input record should answer what entered a training run. An output investigation asks what can explain a particular result. I want both, but I would give them different acceptance criteria.
For the input side, my proposed record includes a dataset version, the source of each collection, the documented basis for its use, and the training runs that consumed it. Where a supplier cannot expose individual records, capture exactly what it can attest to and what remains unknown. A vague assurance should not become a detailed inventory merely because it arrived in a signed PDF.
For the output side, ask what the method actually measures: visual resemblance, a modelled influence estimate, or a counterfactual change after removal. Record the method, the comparison set and the limit of the conclusion. Calling all three “attribution” does not make their answers interchangeable.
This is a proposed review design, not a compliance certification. A well-kept ledger can faithfully record a bad decision. Its job is to make that decision inspectable, not to bless it.
Try the substitution test
As an illustration, imagine a supplier responding to a training-data question with a report that finds no strong match between a generated picture and the images it searched.
I would send that report to the output investigation and leave the input question open. The missing answer is still: which dataset version did this model consume, and what evidence supports the stated basis for using it? Even an excellent answer about this picture cannot supply those missing records.
Now reverse the exchange. A complete dataset manifest arrives in response to a question about a suspicious output. Keep the manifest, but ask for the output investigation too. Knowing the ingredients does not settle every question about the result.
That gives the review a practical failure condition: reject a response that changes the question while appearing to answer it. No new governance platform is required to try this. Put the two questions in separate rows of the next vendor review and refuse to mark one complete using the other's attachment.
The adjacent issue in teacher-lineage governance is preserving the generation chain behind synthetic training data. Here the distinction is narrower: do not replace a record of inputs with an inference from outputs.
Keep the training receipt even when the generated image has no clear attribution; they answer different questions.


