Missing labels pollute the later information environment
Aliases: corpus contamination · unlabeled synthetic media · downstream pollution
What it is
An unlabeled generated street scene is scraped by a news aggregator, written into an entry as “a photograph from the scene,” then swept into the next training run. Downstream actors are not hostile; they ingest under “no mark means a record.” Unlabeled pollution means that unmarked generated artefacts, once in the public information environment, are digested by later search, citation, and training as human records. Error then amplifies as environment, not as a single misread.
Whether one recipient noticed at the time is identifiability. This entry governs how the environment is rewritten after the mark is missing.
Why it happens
Information environments filter on origin class: search engines, newsrooms, dataset pipelines all use “is this a record” as an intake rule. With no mark, the default is almost always “intake as record.” Generated yield outruns human records; unmarked output takes a growing share of later corpora; the next system learns “what the world looks like” from a polluted environment, and synthetic features are treated as world features.
Pollution accumulates. A single misread can be corrected; once a generated artefact is catalogued as “verified material,” correction faces a citation network, not one image.
Studying it
Trace unlabeled generated artefacts into search indexes, encyclopedic citations, open datasets. Measure: times cited as a record, share entering training sets, drift of downstream models on related facts. Independent variables: presence of class metadata, whether the platform detects at intake. Dependent variables: pollution rate (share misclassified as human record), recovery of class after secondary citation.
A single-person lab judgment is not enough. Pollution happens in pipelines and batch intake; methods have to follow the pipeline, not only a questionnaire.
Where it stops holding
Closed, unarchived, no-export sessions barely pollute the public environment. Explicitly fictional communities (role-play, setting bibles) use synthesis as synthesis; missing labels may not cause “treated as record” pollution, though they may still pollute “whose style is this.” Detector false positives that mark human records as generated are reverse pollution; “label more” does not cancel them. This entry does not treat the burden on those who comply, nor that a failed detection is not evidence of human authorship.
Applying it
- Before content enters a state third parties can harvest (public, indexable, downloadable), class must already be in the payload, not only on this site’s chrome.
- Do not let your own search and recommendation pipelines intake unclassed material as human record by default; unclassed should be its own bin.
- When a polluted downstream citation is found, offer a machine-readable correction, not only an edit of the original post.
- Check: pull a sample of published generated artefacts via a public crawler or sitemap; see whether origin class still exists in external indexes. Class at zero outside is an open pollution channel.
Related
- Same group: L3.04.1 Generated content must be identifiable · L3.04.2 Labels must survive after the content circulates
- Nearby: L3.10 Labeling and Watermarking AI Content · L3.03 Hallucination and the Fact-Checking Burden
- Search terms:
corpus contamination·unlabeled synthetic media·information environment