← Blog
AI and societyBusiness models

Business models for the AI internet

By converting website traffic to ad clicks and purchases, websites have historically been able to share their content openly without putting up paywalls.

By Nicholas Wagner ·

By converting website traffic to ad clicks and purchases, websites have historically been able to share their content openly without putting up paywalls. This created an uneasy peace between traffic-driving platforms like Meta and Google and the creators whose work drew people online like writers, journalists, artists, musicians, bloggers, content creators, etc. “Senator, we run ads.”

This business model has societal consequences. Ads incentivize attention. If someone’s video gets a thousand times as many views, they earn roughly a thousand times as much money. Attention can be earned in many ways, some of them less desirable than others.

But AI models are weakening this business relationship. AI applications give users content they want without directing them to the source page, reducing website traffic. This might not be an issue if the websites were compensated for lost ad revenue, but AI models have largely been trained on copyrighted data without paying the copyright holders. In response, many popular websites have restricted access to AI company web crawlers, and users producing valuable data are limiting what they post publicly. This has people like Cloudflare CEO Matthew Prince worrying that the open internet is coming to a close with dire consequences for cultural cohesion.

This post collects emerging business models that might offer a way forward. I focus on models that keep content publicly accessible rather than those that wall it off, such as subscription paywalls or freemium setups that use open citations to drive readers to paid products. Those models work for many, but I’m most interested in approaches that preserve open access and what they incentivize.

Option 1: Pay upfront for licenses

Many major publishers have negotiated deals with AI companies to license their data for training or to be used as contextual references when providing answers. For example, the New York Times, despite suing OpenAI over alleged copyright infringement, has a licensing deal with Amazon where the tech giant pays them at least $20 million annually for the publication’s editorial content. Reddit and StackOverflow, which were both key sources of training data for OpenAI, both have licensing deals with the company.

These arrangements are still in their early days. Both license holders and AI companies are figuring out how valuable specific datasets are to model capabilities. Still, I think the overall picture of what these sorts of deals incentivize is clear. You either need to be such a superstar that an AI app user would revolt if your content did not show up in the generated results, or be part of a large enough bloc of valuable creators to negotiate collectively.

Option 2: Pay-per-crawl

In a pay-per-crawl model, a website owner establishes rules on who can scrape data from their site, how often, and for what price. Cloudflare announced this approach July 1st after building the capability to stop data-hoovering bots that roam the internet.

Pay-per-crawl has some advantages over upfront licensing. It is much easier to automate and scales to small creators who otherwise wouldn’t merit direct deals. AI companies only have so many lawyers to negotiate licensing deals, and they are not incentivized to care about creators with relatively small amounts of data. With pay-per-crawl, anyone can set their price to scrape or be scraped, allowing automatic matching. In a way, it reminds me of how Google Ads democratized access to search advertising.

The model’s weakness is enforcement. As scraping agents become smarter, bad actors will be harder to stop. And just as Google fights SEO spam to keep search results relevant, AI companies will have to figure out which sites in which verticals actually have data worth paying to scrape. I cannot even imagine what the SEO war equivalent looks like in a world where AI companies pay this way.

This model incentivizes content that cannot be easily generated synthetically. News for example concerns things that take place after models have been trained. Even if future models learn continuously over time, they will need some way to access information about current events in the world. Personalized content on special topics is also more valuable than generic “slop” content on mainstream topics. This could offer more upside to creators creating specialized, high-quality, or stylistically distinctive content than large upfront licensing deals.

Option 3: Pay-per-inference

The “Spotify model” compensates music rights owners based on what gets streamed the most. In the AI context, this involves figuring out how much AI companies pay or get paid based on what is generated.

Funnily enough, in the biggest AI music licensing deal to date between ElevenLabs, music licensing organization Merlin, and publisher Kobalt, the artists are not compensated this way. According to Billboard reporter Kristin Robinson, rights holders are paid based on the number of their songs in the licensed dataset and digital proxies such as Spotify streams. That means music that is already popular with human listeners gets compensated more, but those kinds of songs may not be the most valuable when it comes to training ElevenLabs’ model. For example, if most paying customers use Eleven Music to make production music? Or what if certain obscure songs have a disproportionate impact on song quality?

Rights owners could instead be compensated based on the influence of their data on making a piece of AI-generated content. OLMoTrace and related research shows you can trace backwards from a generated output to which training data was most influential in making it. The challenge is scaling these approaches to work with bigger models and getting buy-in from AI companies. If implemented, this could enable proportional royalties and even drive interest in influential sources.

Going back to the music industry, an interesting test case for this just launched in Sweden. Music rights society STIM is partnering with startups Sureel and Songfox to pilot a new license for training AI music models. Royalties flow back up to songwriters based on calculated influence rather than dataset presence and stream counts. I predict these sorts of license arrangements and the SaaS companies that enable them will prove popular with smaller creators who can’t ink traditional licensing deals or believe digital proxies undervalue their work.

I am cautiously optimistic

All of these business models are still getting their tires kicked, but my prediction is that big upfront deals that don’t touch on pay-per-inference royalties will become increasingly rare. The challenge of fairly dividing royalties between groups of rights holders is simply too complex, and lawyers are too expensive for individualized contracts for everyone.

It remains to be seen if these new frameworks can truly make up for lost ad revenue. Still, I am happy that market-oriented solutions to fairly compensate data creators are coming into focus. I won’t claim that optimizing for “maximally influential data” will avoid all the techno-social consequences of optimizing for attention, but at least it offers a new foundation for building a fairer digital economy.


This post was originally published on Learning Journey, our Substack. Subscribe there for new posts on AI and society.

Stay in the loop

Practical AI tips and events, in your inbox.

One really interesting email, plus a note when a new course opens. Unsubscribe anytime.