AI can now generate articles, images, voices, podcasts, videos, influencers, and entire publishing operations at almost zero marginal cost. The question is no longer whether machines can create content. It is what happens when the internet increasingly fills with machine-generated material, future AI systems learn from that material, and people can no longer easily tell who or what created the information they consume.
This page is the canonical research hub for the documentary The Internet Is Eating Itself.
Research summary
The early internet was messy, unreliable, and full of spam. It was also built from millions of original human contributions.
People documented repairs, reviewed restaurants, argued about cameras, edited Wikipedia pages, wrote niche blogs, participated in forums, and shared direct experience. Creating that material required friction. Research took time. Writing took time. Recording and editing took time.
Generative AI changes that cost structure. When another article, image, voice track, or synthetic persona costs almost nothing to create, content production can expand much faster than human attention. The result is not only more low-quality material. It is an information ecosystem in which increasingly polished outputs may draw from the same models, patterns, and source material.
The documentary's core thesis is that the long-term scarcity may not be information. It may be credibility.
Key findings
Publishing friction is disappearing
Generative AI reduces the time and labor needed to produce text, audio, images, and video. This allows one person or organization to create volumes of media that previously required a team.
Content can scale faster than attention
When the marginal cost of another asset approaches zero, the incentive shifts from producing one piece to producing hundreds or thousands.
Synthetic media is becoming less obvious
AI-generated voices, images, and fictional influencers are increasingly plausible to casual viewers. The origin of a piece of media may no longer be obvious from the media itself.
AI tends to reproduce the average
Large language models are built to identify and reproduce patterns. That makes them useful for summarizing consensus and common practice, but it can also make outputs converge around familiar frameworks, opinions, and phrasing.
The training-data loop is changing
AI systems learned from decades of human-created information. As machine-generated material occupies more of the web, future systems may increasingly encounter synthetic material in their data sources. The precise effects remain an active research question.
Trust becomes more valuable as content becomes abundant
When anyone can generate convincing information at scale, audiences must spend more effort evaluating provenance, expertise, experience, and authenticity.
The internet information loop
- Humans create original observations, experiences, arguments, and media.
- Platforms index and distribute that material.
- AI systems learn patterns from large bodies of available information.
- AI systems generate new text, images, audio, video, and synthetic identities.
- Generated content is published back onto the internet.
- Humans and future systems encounter a mixture of original and generated material.
- Determining provenance becomes more difficult.
Research questions
- What happens when machines increasingly create content for other machines to index, summarize, rank, and learn from?
- How much machine-generated material is entering public information systems?
- Can future AI systems reliably distinguish original human evidence from synthetic restatements?
- Does repeated use of similar models cause online language and ideas to converge?
- Which provenance signals remain trustworthy?
- What makes a person, publication, or dataset credible when media can be generated at scale?
Methodology
This project uses the published documentary transcript as its primary source. The dataset separates direct documentary observations, interview-based opinions, interpretive claims, factual claims requiring external support, and future-facing hypotheses.
No uncertain statistic is presented as verified. Quantitative statements, commercial examples, and claims about model behavior are marked in the source ledger when an outside source is required.
People featured
Speaker-level attribution should be checked against the final video edit before individual interview quotes are published as direct quotations.
Platforms mentioned
- Google.
- LinkedIn.
- Wikipedia.
- YouTube.
- Reddit.
- eBaum's World.
Downloads
- Concept dataset
- Claim and verification dataset
- Examples dataset
- People dataset
- Companies and platforms dataset
- JSON dataset
- Source ledger
- Glossary
- Directory
- Methodology
- Information-loop diagram (SVG)
- Full repository: github.com/jessicamalnik/model-collapse-research
Limitations
The documentary includes several statements that were intentionally framed as uncertainty, recollection, prediction, or opinion.
The following require independent sourcing before they are presented as standalone factual claims:
- The exact share or growth rate of AI-generated content online.
- The rate at which future models ingest synthetic content.
- The effects of recursive training on model quality.
- The measurable decline, if any, in search usefulness or review reliability.
Suggested citation
Malnik, Jessica. "Model Collapse: When AI Trains on AI." 2026. https://jessicamalnik.com/resources/model-collapse