In December 2023, The New York Times sued OpenAI and Microsoft, alleging that the companies had used millions of its articles to train models without permission or payment. A few months earlier, OpenAI had gone the opposite direction with other publishers, signing licensing agreements with Axel Springer, the Associated Press, and News Corp that specified what could be used, how, and for what price. The same industry cycle now contains both responses to the same underlying practice.
The shift did not come from a change of heart inside the labs doing the training. It came from a combination of lawsuits like the Times', public backlash, and a growing recognition among content owners that the data being scraped had value the scrapers were not paying for — pressure that made the previously frictionless practice of large-scale scraping considerably more legally and reputationally exposed.
What is replacing the easy-scrape era is a market in negotiated data access. Reddit signed a data-licensing deal with Google in 2024 reported to be worth roughly sixty million dollars a year; Shutterstock struck its own licensing agreement with OpenAI covering its image library. These are actual contracts, with actual payment terms, where there used to be only terms-of-service violations nobody enforced.
The consequence for model developers is that data access is becoming a strategic asset to be secured deliberately, rather than a background resource assumed to be freely available. Whoever locks in exclusive or preferential access to high-quality data — a large archive, a specialized professional corpus, a platform's full historical record — gains an advantage that is harder for competitors to replicate than compute or even talent, because the underlying content genuinely does not exist anywhere else in that form.
This advantages large, well-capitalized labs able to strike expensive licensing deals like the ones Reddit and Shutterstock signed, and disadvantages smaller developers and open efforts that relied on the assumption of freely available scraped data to compete at all. The scraping backlash, in other words, may end up concentrating the industry further rather than distributing accountability more fairly, even though the original complaint was about fairness to data owners.
Content owners are watching each other's strategies closely, because the Times' lawsuit and Axel Springer's license represent two genuinely different bets on the same underlying question: litigate for damages and precedent, or negotiate for ongoing revenue and a working relationship with the companies most likely to shape how information gets accessed going forward. Neither bet has fully paid off yet, and neither publisher can undo its choice now that the other's case is already public.
The scraping era is not fully over — plenty of data collection continues without explicit negotiation, particularly for use cases with less legal exposure than commercial model training. But for the categories of data that matter most, the direction is being set in a Manhattan federal courtroom and in Reddit's boardroom at the same time, and every publisher deciding which path to take is watching both outcomes before committing to either.
