Explainer · From Session 009

The Era of Free Training Data Is Over

Reddit walled its garden and vanished from ChatGPT in a day. The web's content is becoming inventory, and yours is part of it.

The short answer

Training data is the content an AI model learns from, and the era of getting it free is ending. Reddit is walling off its content to force AI companies to pay, and its citations inside ChatGPT collapsed on August 14, 2026. For business owners, the shift makes owned audiences and first-party content more valuable, not less.

Key takeaways

What is training data, and why does Reddit want to be paid for it?

Training data is the text, images, and code an AI model learns from before anyone ever prompts it. Model quality tracks data quality, which makes two decades of human conversation on Reddit one of the most valuable text corpora ever assembled, and until recently AI companies absorbed it for free.

Reddit noticed. Justin Novak's read on Session 009: Reddit is walling off its content to force whoever wants to train on their data to pay for it, because nobody gets to endlessly scrape decades of rich content anymore. His conclusion in a sentence: the era of free training data is ending. As he added, commercializing the data instead of donating it makes total sense from Reddit's side of the table.

What changed in August 2026?

The wall became visible on August 14, 2026, when Reddit citations in ChatGPT fell off a cliff, dropping from the number one cited source toward zero in a day. The question that opened the discussion said Reddit might block Google entirely because AI Overviews killed their traffic. The full marketing fallout is in our companion post on the Reddit citations collapse.

Paywalls for machines are spreading beyond Reddit. Justin described another company blocking all bot traffic unless scrapers pay in microtransactions per page view. Publicly, Cloudflare now offers exactly that model, a pay-per-crawl marketplace where AI crawlers pay publishers per page fetched. Content is becoming inventory, with a price per unit.

“Reddit is walling off content to essentially force whoever wants to train on their data to pay for it. You can't just endlessly scrape their decades of really rich content.”

Justin Novak · 14:20

Who wins and who loses when data gets walled off?

Platforms holding unique data win. Reddit, publishers, and any site with content machines want can now charge rent on what it used to give away. AI companies lose a free input and gain a bill, which they can afford.

Businesses that rented their audience lose quietly. The companies paying Reddit marketing agencies to earn LLM citations watched the mechanism disappear overnight, without violating a rule or missing a payment. Ian Kilpatrick's warning covers them exactly: platforms make fickle changes, and suddenly your entire business model is out the door. His prescription is the old one because it works: multiple eggs in multiple baskets.

What should a business owner do about it?

Own the assets that no platform can reprice. An email list, a website on your own domain, and first-party content are the marketing equivalents of holding the deed instead of renting. Every wall that goes up around third-party data makes owned channels relatively more valuable.

Treat your own content as the appreciating asset it just became. In a world where models pay for quality text, the transcripts, case studies, and proprietary numbers only you can publish are exactly what AI systems still want and cannot get elsewhere. AI4NTP (AI for Non-Techy People) runs on this playbook: every session becomes owned transcripts, tools, and posts on our own domain, and the AI assistants cite us for facts nobody else has.

Diversify the rest. The Reddit cliff is one platform repricing one input, and it will not be the last. Owners who spread across channels and keep the core assets in-house get to watch these fights as spectators instead of casualties.

“I've seen Google do this, I've seen so many platforms make a fickle change, and suddenly your entire business model's out the door.”

Ian Kilpatrick · 9:48

Frequently asked questions

Why is Reddit charging AI companies for its data?

Two decades of human conversation is one of the most valuable training corpora in existence, and Reddit was giving it away to scrapers. Walling it off and forcing AI companies to pay converts that archive into revenue, which, as Justin Novak put it on Session 009, totally makes sense from a business point of view for Reddit.

What is pay-per-crawl?

A model where a site blocks bot traffic unless the scraper pays per page viewed. Justin Novak described a company doing this with microtransactions on Session 009, and Cloudflare publicly offers the same model, letting publishers charge AI crawlers per page fetched.

Does the end of free training data affect my website?

Directionally, yes, and in your favor. Your first-party content becomes more valuable as freely scraped data disappears, both as something AI systems cite and as an asset you control. The businesses hurt by the shift were the ones renting visibility on platforms like Reddit rather than owning their channels.

How do I protect my business from platform changes like this?

Ian Kilpatrick's rule: multiple eggs in multiple baskets. Diversify your channels, and keep the core assets where no platform can reprice them, meaning your email list, your own domain, and content only you can produce. One fickle change should never be able to take your pipeline to zero.

Tools used in this post

Every tool here has its own page with pricing, who used it live, and honest alternatives.

Where this came from

Session 009 · Recorded live
AI office hours: your questions, answered live.
Watch the recap →YouTube replay →
7:52 JustinI read Reddit might block Google entirely because AI Overviews killed their traffic. As a business owner, should I still be investing in SEO?
14:20 JustinReddit is walling off content to essentially force whoever wants to train on their data to pay for it. You can't just endlessly scrape their decades of really rich content. So I think the era of free training data is ending.
14:20 JustinI've even seen another company implement something similar, where they block all bot traffic, or if you do want to scrape it, you pay in microtransactions for every page view.
14:20 Justinthem trying to commercialize their data instead of it just being scraped for free, which totally makes sense from a business point of view for Reddit.
11:51 Justinwhat this graph is showing is Reddit citations falling off a cliff here on August 14th.
9:48 IanI've seen Google do this, I've seen so many platforms make a fickle change, and suddenly your entire business model's out the door.
9:48 IanYou have to have multiple eggs in multiple baskets.

Who wrote this

Every AI4NTP post is written by an operator who was in the room when the work happened.

Justin Novak
Partner at AI4NTP

Founder and host of AI4NTP. He sold his first company from his college dorm room, and as a fractional CMO has helped scale multiple businesses past $50M in ARR.

Ian Kilpatrick
Partner at AI4NTP

A designer, developer, and serial entrepreneur writing code since age 10. He has worked with Disney, the Golden Globes, and the AMAs, and now runs a fleet of AI agents doing real work every day.

Alec Saluga
Partner at AI4NTP

A former B2B salesman with no technical background who self-taught AI. He builds and deploys AI-driven marketing and websites, and has grown a following of over 15,000 teaching AI adoption.

More field notes

Want to watch this happen live?

AI4NTP runs free live sessions where real operators build in real time, with the audience picking what gets built. No theory, no slides.

← All posts Get these in our newsletter
Last updated 2026-08-20