After Computing Power Comes Data: Is AI's New Bottleneck Reshaping Pricing Power in the Content Industry?

Mô hình/nhà cung cấp liên quan: Anthropic Nhà cung cấp
After Computing Power Comes Data: Is AI's New Bottleneck Reshaping Pricing Power in the Content Industry?

In August 2026, A-share AI data-related stocks surged collectively. Dook Culture hit the 20cm daily limit, while copyright content stocks like CITIC Press and China Publishing followed suit, and data service provider Hithink RoyalFlush also strengthened. Capital is chasing a newly emerging logic: the fuel for AI large models is running low. According to research institutions like Epoch AI, high-quality public text data may face significant shortages from the mid-to-late 2020s to the early 2030s. Liu Liehong, director of the National Data Administration, made a key point at the 2026 World Intelligence Industry Expo: "AI is like a data refinery; competitiveness never depends on the size of the furnace, but on the grade of the ore."

Computing power follows Moore's Law, but data does not. As large model parameter scales surge from hundreds of billions to trillions, the volume of training data they consume is expanding exponentially. For text models, GPT-2's training data was on the order of tens of gigabytes, while GPT-3 used hundreds of gigabytes. Moreover, large models are evolving from pure text to multimodal. According to information disclosed at the 2026 Robot Full Industry Chain Conference, achieving practical embodied intelligence requires at least 10 million hours of multimodal data, and considering multi-scenario adaptation, multimodal types, data yield, and other factors, the actual scale needed may be even higher. According to public data from the Guizhou Provincial Big Data Bureau, currently compliant data for real physical interaction scenarios in China is only 500,000 hours, with a gap exceeding 99%. The natural pasture on the internet is being consumed, and a risk often amplified by public opinion is that AI-generated content's data self-pollution may accelerate the tightening of high-quality original text supply.

Article image

I. The Boomerang Effect of AI-Generated Content

When AI-generated content floods the internet and is then scraped by next-generation AI models for training, what happens? An increasingly thorny issue is model collapse—if models are iteratively trained on second-hand AI-generated data, they gradually deviate from the real-world distribution and lose the ability to understand rare events. This is like a closed photocopying system: copying copies loses information with each generation. However, academia remains divided on the severity of model collapse—only indiscriminate, large-scale training on pure AI-generated data triggers severe collapse. In practice, multiple mitigation strategies exist: strict filtering, data cleaning, deduplication, and mixing with original human data can effectively suppress degradation. AI content is not inherently toxic; the key lies in governance.

To obtain original corpora uncontaminated by AI-generated content, Silicon Valley tech giants have begun mining the value of traditional media manuscripts and paper books. According to reports from The Washington Post and other media, Anthropic launched a book digitization project codenamed "Project Panama" in 2024. Internal documents stated: "Project Panama refers to our destructive scanning of all books in the world... We do not want the outside world to know we are doing this." Previously, the company was sued by author groups for downloading over 7 million e-books from pirated resource libraries like LibGen for model training. On July 20, 2026, Judge Araceli Martínez-Olguín of the U.S. District Court for the Northern District of California signed a final approval order approving Anthropic's $1.5 billion settlement for copyright infringement claims, covering approximately 482,000 to 500,000 works, at about $3,000 per work. This is the largest settlement in U.S. copyright history. In the case, retired Judge William Alsup's summary judgment opinion suggested that scanning legally purchased books for internal model training has transformative use characteristics and falls under fair use. However, two points must be emphasized: first, the ruling was only a pre-trial district court opinion; the case ultimately settled without appeal and does not constitute binding precedent nationwide in the U.S.; second, the opinion applies to training on legally obtained books and cannot be extrapolated to justify destructive book scanning as commercially legitimate. Additionally, the entire reasoning is rooted in the U.S. copyright system and cannot be directly transplanted to China's copyright environment.

Article image
Article image

II. Traditional Content Providers' New Role: AI Miners

When original corpora shift from a free lunch to a scarce mineral, the value chain is undergoing a silent restructuring. But it must be clarified that AI companies purchase text for two distinctly different purposes: one is for model pre-training or SFT fine-tuning, and the other is solely for RAG (Retrieval-Augmented Generation) knowledge base retrieval after deployment. The legal authorization scope, commercial pricing, and copyright risks differ vastly between the two. Currently, many AI companies purchase copyrighted text only for retrieval, not training, so one cannot simply expect all copyrighted content to sell at high training-data prices.

AI companies are beginning to pay for high-quality corpora. HarperCollins reached an AI training licensing agreement with Microsoft for selected non-fiction books, limited to some older titles, with a total price of $5,000 per book for a three-year license, split evenly between publisher and author, on an opt-in basis, with a cap on text citation ratios in the contract. Taylor & Francis signed a $10 million academic content licensing agreement. According to MarketIntelo forecasts, the global training data licensing market will reach $22.6 billion by 2034 (the statistical scope covers multiple types of training datasets, not limited to text copyright content). However, caution is needed regarding the reference significance of overseas cases: some licensing deals come with copyright litigation settlement backgrounds, and prices do not represent normalized market pricing.

Domestically, progress is also accelerating. Multiple sources close to publishing houses and image copyright agencies revealed that some AI companies have begun purchasing compliant datasets from traditional content institutions. The National Data Administration has explicitly proposed building a dataset value system based on tokens. But all this is still in its early stages. Between owning copyright and becoming an AI miner lie significant engineering and costs: digitization, cleaning, annotation, desensitization, format conversion, and more. Chinese corpora face not only complex copyright ownership issues but also practical constraints such as data dispersion, immature circulation mechanisms, insufficient deep annotation, and vertical processing capabilities. More importantly, not all copyright libraries can enjoy value revaluation—most popular general text has very low per-token value, and only professional, exclusive, scarce vertical content may command high premiums. A research report from Industrial Securities judges that AI is pushing the media content industry from consumer goods to dual pricing of "data + IP assets." But this judgment requires the supporting systems of rights confirmation, trading, and traceability to mature simultaneously.

Article image
Article image

III. Alternative Paths: AI Companies Can Survive Without Grabbing Books

If corpus shortages truly become a hard constraint, AI companies will not sit idly waiting for publishers to mine. The potential of synthetic data is being validated. Studies have shown that high-quality synthetic data can partially replace real data in fields like mathematics and code, significantly reducing reliance on original human text. Improvements in data efficiency are also changing the accounting logic. Model architecture improvements (such as MoE) and training method optimizations (such as curriculum learning) are reducing the amount of raw data needed per unit of intelligence. RAG and continual learning greatly expand the boundaries of information access—large models do not need to memorize all knowledge during pre-training but can retrieve on demand during inference.

Additionally, AI companies have multiple alternative paths: building in-house manual annotation teams to produce vertical data, purchasing from professional data service providers (such as Hithink RoyalFlush), scraping open compliant public data, procuring overseas multilingual datasets, and directly signing authors rather than going through publishers. Traditional book copyrights are just one data source, not irreplaceable. High-quality public text in Chinese is relatively scattered, and data circulation mechanisms for professional publishing, academic, and industry reports are immature. Much data exists but is difficult to circulate compliantly. Publishing institutions have content assets but may lack data engineering capabilities.

IV. Three Barriers: From Ore Vein to Pricing Power

Even if corpus scarcity holds, there are three practical barriers between owning a copyright library and mastering pricing power.

First, copyright ownership is complex. Publishers mostly hold only publishing format rights and information network transmission rights; text copyright belongs to authors. To license for AI training, negotiations with each author are required; many old books have lost authors or broken authorization chains. According to industry insiders, a leading publisher attempted to license its books to an AI company, but because many author contracts did not include AI training rights, the number of titles ultimately licensed was far below expectations, with negotiations lasting months.

Article image
Article image

For publishing institutions to truly master pricing power, they cannot rely solely on existing copyrights; they also need capabilities in copyright confirmation, data cleaning and processing, compliant licensing, and secure circulation. Existing ancient books and old books can supplement the corpus pool in the short term, but long-term sustainable clean original corpora depend on continuous new human creation. Existing copyright libraries can only serve as a buffer, not permanently solve data iteration needs.

Second, the token-based pricing system is not yet mature. How to count how many tokens a book contributes to training, how to trace, how to share revenue, and how to prevent data theft and secondary circulation all lack technical and regulatory support. Currently, overseas deals mostly use package licensing (annual fees, whole-library purchases), and precise per-token trading models are still far from implementation.

Third, AI companies' alternative supply is taking shape. Synthetic data, data service providers, open-source data, and direct author signings—traditional copyright holders' bargaining power is being diluted by multiple paths simultaneously. Additionally, a large number of ancient books, modern public domain texts, and open academic resources are permanently free, continuously suppressing the corpus pricing ceiling for ordinary popular books. In the future, if China introduces a statutory licensing mechanism for AI training data (uniform rates, no one-on-one negotiations), it will completely change the current game—AI companies would not need to obtain author authorization one by one, and publishers' exclusive bargaining power would be greatly weakened.

Article image

A more accurate description of the corpus bottleneck might be layered scarcity: what will truly be scarce in the future may not be ordinary public text, but high-quality, exclusive, compliant, traceable, professional vertical data, as well as multimodal and interactive data. For example, medical textbooks, legal precedents, and in-depth vertical industry reports are more likely to form scarce supply; while large volumes of homogeneous online popular novels, public domain ancient books, and general encyclopedia entries have relatively abundant supply.

Article image

Returning to the A-share market in August, the surge in the data sector reflects capital's hunger for the new narrative of AI corpora, not performance realization. Before the rights confirmation system, pricing rules, and revenue-sharing mechanisms are fully operational, publishing companies' AI corpus revenue is likely to remain at the level of intent agreements. In the long run, as AI competition continues to evolve toward multimodal, demand for image, video, and robot interaction datasets will keep rising, and the strategic weight of pure text copyright corpora will gradually be diluted. The so-called corpus scarcity is not limited to text materials.

Industry potential solutions include collective management organizations for copyright negotiating uniformly and mandatory revenue-sharing mechanisms for AI licensing. If long-term revenue concentrates in platforms and publishing institutions while frontline creators struggle to benefit, the motivation for high-quality original content creation will continue to decline. Clean, compliant, exclusive original corpora will become important competitive chips. But simply holding existing text does not equal a moat. Corpus scarcity does not necessarily give traditional publishers pricing power; pricing power is a dynamic result of multi-party games. What truly determines bargaining power is: who can continuously produce high-quality original content, who can transform copyright assets into compliant, traceable, tradable data assets, and who can establish stable revenue-sharing mechanisms among AI companies, platforms, and authors. Merely hoarding existing text cannot build a sustainable competitive barrier in the AI era.

Chia sẻ bài viết này