Why AI is Running Out of Data—and What It Means for Your Business

Explore the AI data shortage crisis in 2025. Learn how data exhaustion affects e-commerce, synthetic data risks, and strategies to stay competitive.

Rihards Ručevics23 min read
Why AI is Running Out of Data—and What It Means for Your Business
Why AI is Running Out of Data—and What It Means for Your Business

Introduction: The AI data crisis is here in 2025

At Pickastor, our analysis shows that the question is no longer whether AI is running out of data, but how quickly the shortage will reshape competitive advantage across every digital industry, including e-commerce.

Authentic training data potentially depleted within ~6 years A Stanford-linked report cited by Forbes warns that the **pool of authentic data available for training AI models could be depleted within six years**, highlighting fears that scaling laws based on ever larger datasets may soon hit a ceiling. Stanford report via Forbes (2026)
High-quality text exhausted before 2026; low-quality text 2030–2050; vision data 2030–2060 Original 2022 projections cited in 2024 analyses forecast that **high‑quality language data** for AI training would be exhausted **before 2026**, with **low‑quality language data** running out between **2030 and 2050** and **vision (image) data** between **2030 and 2060** if current trends continue. Epoch AI projections summarized by NextBigFuture (2024)

The supply of quality training data is shrinking faster than expected

The numbers are striking. According to PBS NewsHour (2024), the global stock of high-quality, human-written text suitable for large language model training could be exhausted as early as 2026. Epoch AI places the total usable supply at roughly 15 to 20 trillion tokens, with a median exhaustion forecast of 2028 and an outer boundary of 2032. That is a remarkably narrow window given how aggressively AI capabilities are expected to expand in the same period.

This is not a theoretical concern confined to research papers. According to Forbes (2026), a Stanford-linked report warns that the pool of authentic data available for AI training could be depleted within six years. Elon Musk stated publicly in early 2025 that AI firms have already exhausted the cumulative sum of human knowledge available for training purposes, a claim that, while contested in its framing, reflects a genuine and growing consensus among AI researchers and industry leaders.

What this means for e-commerce businesses specifically

For SMB owners, enterprise teams, and marketplace sellers, the implications are direct and practical. AI tools that power product discovery, search ranking, recommendation engines, and content generation all depend on continuous access to fresh, high-quality data. As the broader supply of public human-written text tightens, the authenticity and relevance of that underlying data becomes a critical differentiator.

Businesses that understand the data shortage timeline now are better positioned to adapt their content strategies, optimize how their product information feeds into AI systems, and maintain visibility as these systems grow increasingly selective about the data they can effectively use.

The sections that follow break down exactly how this crisis is unfolding, trend by trend, and what your business can do about it.

Trend 1: High-quality text data exhaustion accelerating toward 2026–2028

The timeline for AI training data exhaustion is arriving faster than most organizations anticipated. Early projections from 2022 suggested the problem would emerge before 2026, but revised forecasts now place the critical window between 2026 and 2032, with the median estimate converging around 2028 as the most likely inflection point for large-scale model development.

High-quality text data exhaustion 2028 year (median forecast)
Global stock of training-suitable text 17.5 trillion tokens
Low-quality text data exhaustion window 2040 year (midpoint)
Vision/image data exhaustion window 2045 year (midpoint)
15–20 trillion tokens; exhaustion window 2026–2032; median 2028 Epoch AI estimates the global stock of high-quality public human-written text suitable for large language model training is around **15–20 trillion tokens**, with exhaustion of this supply likely **between 2026 and 2032** and a **median forecast of 2028**. Epoch AI via AIToolDiscovery (2024)

How the forecasts have shifted

When researchers first modeled the supply of high-quality, human-generated text on the public internet, the concern was largely theoretical. By 2024, it had become a concrete operational problem. According to PBS NewsHour (2024), Epoch AI's widely cited analysis estimates that tech companies could exhaust the available stock of public human-written text data for chatbot training somewhere between 2026 and 2032. That is not a distant horizon. For businesses building strategies around AI-driven discovery and visibility, 2026 is effectively the next planning cycle.

The upward revision from "before 2026" to "2026 to 2032" might sound reassuring, but it reflects a more nuanced problem: the data that remains is increasingly lower in quality, narrower in scope, or already heavily used across multiple training runs. The quantity of available text has not grown proportionally with the appetite of modern large language models.

Why 2028 matters as a median forecast

The 2028 median estimate is significant because it represents the point at which most major AI labs are projected to face genuine constraints on scaling through conventional data acquisition. Before that point, organizations with access to proprietary, high-quality, structured data will have a measurable advantage in training more capable and accurate models.

IBM research reinforces the urgency of the near-term window, projecting that publicly available human-generated data could be effectively exhausted as early as the end of 2026. That projection compresses the timeline considerably and explains why AI companies are accelerating proprietary data licensing deals at pace.

For e-commerce businesses, this trend carries a direct implication. The data your business generates, including product descriptions, customer reviews, and structured catalog content, is precisely the kind of high-quality, domain-specific text that becomes more valuable as general public data grows scarcer. Understanding this dynamic now positions your business to act before the window narrows further.

Trend 2: Synthetic data becomes the default, but quality risks mount

With high-quality human text growing scarce, AI labs have turned to an increasingly common workaround: generating training data using AI itself. Synthetic data now fills gaps that real-world content can no longer cover, and its use is accelerating across every major model development pipeline.

Anthropic forecast: high-quality text exhaustion 2027 year
IBM Research: publicly available data exhaustion 2026 year (end of)

The synthetic data feedback loop

The appeal of synthetic data is obvious. It is scalable, cost-effective, and can be produced on demand. The problem is what happens when models begin training on their own outputs at scale. Researchers describe this as "model collapse," a gradual degradation process where each successive generation of AI learns from slightly distorted versions of the previous one. Errors compound. Nuance erodes. The model drifts further from the grounded, human-generated knowledge it was originally built on.

According to SAP's GenAI analysis (2024), this feedback loop represents one of the most underappreciated structural risks in modern AI development. The concern is not hypothetical. Major analyses in 2025 and 2026 have shown that mixing even 0.1% low-quality synthetic data into a training set can irreversibly degrade model performance in certain configurations. That threshold is remarkably low, and it signals that quality control in synthetic pipelines is far harder than it appears.

Why e-commerce is especially exposed

For e-commerce teams, the synthetic data problem is not abstract. It arrives in your product catalog, your review summaries, and your AI-generated copy. As more brands adopt AI content tools to scale their listings, a growing proportion of what those tools produce is itself trained on synthetic outputs. The result is content that sounds plausible but lacks the specificity, accuracy, and authentic voice that converts browsers into buyers.

Consumer trust is measurable. Synthetic product descriptions and AI-generated reviews contribute to what researchers are calling "slop fatigue," a growing audience resistance to content that feels generic or machine-produced. Conversion rates suffer when shoppers sense that product information lacks genuine authority.

Understanding how data annotation actually works in practice helps clarify why human-generated, domain-specific content holds such a structural advantage over synthetic alternatives. Annotators make judgment calls that AI cannot reliably replicate, and those judgment calls are precisely what makes training data useful.

The shift from more data to better data is not a philosophical preference. It is becoming an operational necessity.

Trend 3: Content licensing and data access tightening across platforms

The "better data" imperative is colliding with a hard structural reality: the platforms that host the most valuable human-generated content are actively closing the door on AI developers. Licensing restrictions and API policy changes are shrinking the pool of freely usable, high-quality web text at precisely the moment demand for it is highest.

Major platforms restrict unauthorized AI training access

According to PBS NewsHour (2023), the AI industry's gold rush for chatbot training data is accelerating the depletion of usable human-written text. Platforms such as Reddit and Stack Overflow have responded by tightening commercial API terms, explicitly prohibiting scraping for AI model training purposes. Major news organizations have followed, strengthening licensing agreements to prevent their editorial content from being used without compensation.

The practical effect is significant. Large volumes of structured, domain-specific, human-authored content that AI developers previously accessed freely now sit behind commercial agreements or are simply unavailable. The open web, once treated as an unlimited resource, is becoming a governed one.

AI vendors pivot toward proprietary data deals

With public web scraping becoming legally and commercially untenable, AI vendors are being pushed toward a fundamentally different model. Proprietary data deals, enterprise partnerships, and licensing negotiations are replacing bulk crawling as the primary route to quality training material. This raises costs and creates structural advantages for organizations that already hold large, well-organized datasets.

For e-commerce businesses, this shift carries a direct implication. Marketplaces and platforms are beginning to restrict data access in ways that mirror what happened in the broader web content space. First-party product data, including catalog information, customer interaction records, and transaction histories, is becoming a genuinely scarce and strategically valuable asset. Understanding how data shapes AI-driven analysis makes clear why proprietary data ownership is moving from a technical concern to a boardroom priority.

Businesses that treat their own data as a commodity risk ceding a structural advantage to competitors who recognize its growing scarcity value.

Trend 4: Data infrastructure and governance emerge as competitive advantages

The strategic response to AI running out of data is not simply to collect more of it. Forward-looking organizations are shifting their focus from data volume to data quality, investing in the infrastructure and governance frameworks that make their existing first-party data genuinely useful for AI systems.

A split-screen diagram comparing a disorganized data warehouse with scattered, unlabeled nodes on the left versus a clean, structured canonical data model with color-coded validation layers and provenance tracking on the right

From "more data" to "better data"

According to SAP's GenAI analysis (2024), the data scarcity problem is fundamentally an infrastructure problem. Companies that built their AI strategies around harvesting vast quantities of public web data are now finding that source increasingly restricted, legally contested, and of declining quality. The organizations gaining ground are those that invested earlier in canonical data models, validation frameworks, and systematic drift monitoring to ensure their proprietary data remains accurate, consistent, and AI-ready.

This is a meaningful strategic shift. Data governance, provenance tracking, and quality assurance were historically treated as IT housekeeping tasks. They are now competitive differentiators. Knowing where your data came from, how it has changed over time, and whether it still accurately reflects reality is the kind of operational discipline that separates AI-capable businesses from those that are not.

Structured first-party data as a moat

For e-commerce teams specifically, this trend has direct and immediate implications. As public data becomes scarcer and more contested, AI systems will increasingly favor structured, well-attributed, first-party sources. Businesses with clean product catalogs, complete attribute coverage, and consistent structured data will be more visible to AI-driven discovery and recommendation systems than competitors relying on incomplete or poorly formatted listings.

Catalog hygiene is no longer a back-office concern. Attribute completeness, schema consistency, and structured data coverage are now front-line competitive priorities. It is also worth understanding how platforms like Notion handle data ownership, since the tools businesses use to manage information may themselves have implications for data control and provenance.

Organizations that treat their internal data infrastructure as a strategic asset, rather than a technical overhead, will be structurally better positioned as AI scarcity intensifies.

Trend 5: E-commerce pivot to structured, AI-readable product data

As generative models face a data crunch, e-commerce teams are responding with a sharply practical priority: making their product data legible to AI systems. Structured schemas, clean feeds, and rigorous catalog hygiene are no longer technical niceties. They are visibility strategies.

Years until high-quality text exhaustion (Stanford estimate) 6 years from 2024

Why AI systems depend on structured product data

AI search systems and AI Overviews do not browse product pages the way a human shopper does. They parse structured signals: schema markup, attribute completeness, freshness of feed data, and review metadata. When global training data growth slows, these systems become more selective, not less. Products that surface reliably are those with clean, current, and properly formatted data behind them.

The implication is direct. An AI Overview recommending running shoes will favor listings with complete Product schema, verified Review markup, and up-to-date pricing over listings with rich prose descriptions but no structured signals. The quality of the writing becomes secondary to the quality of the data architecture.

The invisibility risk for unstructured catalogs

Unstructured product descriptions and incomplete attribute sets are increasingly invisible to AI-driven discovery. This is not a future risk. It is an active one. As researchers warn we could run out of data to train AIs by 2026, AI systems will increasingly rely on the structured, machine-readable data that remains available rather than attempting to extract meaning from unstructured text at scale.

For marketplace sellers, this creates a compounding disadvantage. Incomplete product attributes, missing GTINs, and stale inventory feeds all reduce the probability that an AI system will surface a listing with confidence. The gap between structured and unstructured catalogs will widen as data scarcity intensifies.

Investing in structured data now as a long-term advantage

Sellers and store owners who invest in structured data infrastructure today are building a durable competitive position. This includes implementing FAQPage and Review schemas, maintaining fresh product feeds, and auditing catalog completeness regularly. Understanding how data science intersects with AI-driven systems is increasingly relevant for e-commerce teams making these decisions.

The businesses that treat catalog data quality as a strategic function, rather than a back-office task, will maintain AI visibility precisely when competitors with unstructured data become harder for AI systems to find and recommend.

What this means for your e-commerce business in 2025

The macro-level AI data shortage is not an abstract research problem. For e-commerce businesses operating in 2025, it translates directly into concrete risks around product discoverability, marketplace ranking, and customer conversion. As AI systems grow more selective about the data they trust and surface, the gap between well-optimized and poorly-optimized product catalogs will widen significantly.

AI visibility is no longer automatic

AI-powered search engines, shopping assistants, and recommendation engines no longer surface products simply because they exist in a catalog. These systems actively evaluate data quality, attribute completeness, and structural consistency before deciding what to recommend. Products that lack structured markup, complete specifications, or clear category taxonomy are increasingly invisible to AI retrieval layers, regardless of how strong the underlying product actually is.

Optimizing for AI discovery now requires the same deliberate attention that SEO received a decade ago. Understanding AI readiness at the catalog level is the starting point for any business that wants to remain visible as AI-mediated shopping continues to grow.

The synthetic content trap

Many e-commerce teams have turned to AI-generated product descriptions, bulk-generated attributes, and templated review content to scale their catalogs quickly. This approach is becoming a liability. As SAP notes in its analysis of generative AI data quality, AI systems trained on synthetic or low-quality data produce degraded outputs, and the same logic applies in reverse: AI retrieval systems are increasingly capable of identifying and downweighting thin, synthetic, or repetitive product content. Stores relying heavily on generated content risk ranking penalties and reduced recommendation frequency precisely when AI-driven traffic is accelerating.

First-party data quality as a competitive moat

Stores with clean, human-verified, richly attributed product data hold a structural advantage that compounds over time. In our experience at Pickastor, the catalogs that perform best in AI-driven environments share three characteristics: complete attribute coverage across every SKU, consistent taxonomy applied at scale, and regular data audits that catch gaps before they affect visibility. This is not a one-time project. It is an ongoing operational discipline.

Heightened risk for marketplace sellers

Marketplace sellers face a specific and growing vulnerability. Platforms including major retail aggregators are tightening data access policies, limiting how sellers can structure and export their own product information. This makes owning and optimizing your independent product feed critical. Sellers who rely entirely on marketplace-managed listings have limited control over how AI systems interpret and rank their products, making off-platform catalog hygiene a non-negotiable priority for 2025.

Predictions and outlook: What to expect beyond 2025

The data scarcity problem is not a distant theoretical risk. The timelines are specific, the forecasts are converging, and the structural shifts they imply for AI development and e-commerce strategy are already underway. Understanding what the next three to five years look like gives businesses a meaningful window to act before competitive advantages close.

High-quality public text will effectively run out by 2028

According to PBS NewsHour (2024), the supply of human-written text suitable for large language model training could be exhausted as early as 2026. Epoch AI places the global stock of high-quality public text at roughly 15 to 20 trillion tokens, with a median exhaustion forecast of 2028 and an outer boundary of 2032. Lower-quality text and image data have longer runways, projected to deplete somewhere between 2030 and 2060, but the premium training material that drives model quality improvements will be gone first.

This matters because AI labs have built their competitive moats on scale. Once the public internet is effectively mined out, the next frontier is proprietary enterprise data. Expect aggressive licensing deals, exclusive data partnerships, and new monetization models from platforms that control large structured datasets.

AI scaling will slow without new data sources

Model performance improvements have historically tracked closely with training data volume. As high-quality public text becomes scarce, AI labs are pivoting toward video, audio, and synthetic data generation. However, synthetic data quality remains inconsistent, and its ability to fully substitute for authentic human-generated content is unproven at scale. This creates a period of likely deceleration in frontier model capabilities, at least until new data pipelines mature.

The financial pressure compounds the technical challenge. Forecasts synthesized by analysts highlight a potential $800 billion shortfall by 2028 in revenue needed to fund AI's projected computing infrastructure demand, a gap closely tied to data scarcity and data integrity constraints. Investment in AI infrastructure may outpace the data supply needed to justify it.

Owned product data becomes a premium business asset

For e-commerce operators, the most actionable implication is straightforward: structured, high-quality product data will carry increasing commercial value. Platforms will charge for data access or require exclusivity. Businesses that have invested in clean catalogs, structured markup, and rich product attributes now will hold a durable pricing and visibility advantage as AI systems compete for reliable training and retrieval inputs.

The window to build that advantage is narrowing. Stores that treat product data as infrastructure today are positioning themselves ahead of a market where data quality is no longer a technical detail but a core competitive differentiator.

Year-over-year comparison: How the data crisis narrative shifted in 2025

The story of AI's data problem did not emerge fully formed. It evolved, and the shift between 2024 and 2025 marks one of the most significant turning points in how the industry understands and responds to this constraint. What was once a theoretical concern has become an operational reality.

A split-screen timeline graphic showing two columns labeled 2024 and 2025, with contrasting status indicators moving from amber warning signals to red alert markers across three rows: data scarcity, synthetic data risks, and e-commerce data quality

From debate to acknowledgment: The data scarcity timeline

In 2024, the question of whether AI was running out of training data was largely academic. Researchers published projections, timelines varied widely, and most major labs treated the issue as a future problem to be solved by future innovation. The debate centered on when scarcity would arrive, not whether it already had.

By 2025, that framing collapsed. According to The Guardian (2025), Elon Musk stated publicly that AI firms have already exhausted all available human data for training purposes. Neema Raphael, Chief Data Officer and Head of Data Engineering at Goldman Sachs, made a similarly direct assessment, stating that the industry has already run out of data when referring specifically to the shortage of suitable training inputs for frontier models. These are not cautionary projections. They are present-tense declarations from people operating at the center of the problem.

Synthetic data: From solution to cautionary tale

In 2024, synthetic data was widely positioned as the answer to training data scarcity. The logic was straightforward: if real-world data is finite, generate more. Pilot programs were promising, and enthusiasm was high.

By 2025, that optimism had become more measured. Researchers and practitioners began documenting the risks in concrete terms, including model collapse from feedback loops, quality degradation when synthetic outputs re-enter training pipelines, and the compounding effect of errors across model generations. Synthetic data remains part of the toolkit, but it is no longer treated as a clean solution.

E-commerce data quality: From optimization to obligation

Perhaps the most consequential shift for e-commerce operators is the change in how product data quality is perceived. In 2024, investing in structured attributes, clean catalogs, and rich product metadata was considered a best practice, a way to improve performance at the margins.

In 2025, it is a competitive necessity. As AI systems increasingly determine search visibility, marketplace ranking, and recommendation placement, the quality of a business's underlying data directly shapes its commercial outcomes. Optimization has become infrastructure.

Expert roundup: What industry leaders are saying about AI data exhaustion

The consensus forming among researchers, executives, and technologists is striking in its urgency. Across institutions and industries, prominent voices are converging on a shared conclusion: the data that powered the first generation of large language models is running out, and the implications extend well beyond the AI labs themselves.

Epoch AI and Anthropic: The research community sets a timeline

Researchers have been the most precise in framing the problem. Epoch AI estimates that the supply of clean, high-quality, human-written public text suitable for large language model training will likely be exhausted between 2026 and 2032, with 2028 as the median exhaustion date. That is not a distant horizon. For businesses making infrastructure decisions today, it represents a window of two to six years before the data pipeline that feeds frontier AI fundamentally changes.

Anthropic researchers have offered an even tighter projection. At current model training growth rates, according to PBS NewsHour (2024), all publicly available high-quality text could be exhausted by 2027. That implies a near-term ceiling on scaling laws built primarily around feeding models more data.

Goldman Sachs: Scarcity is already here

Industry leaders are not waiting for 2027 to sound the alarm. Goldman Sachs' Chief Data Officer, Raphael, has stated plainly that suitable training data for frontier AI models is already scarce. "We've already run out of data," Raphael said, referring specifically to the shortage of high-quality inputs despite the vast volume of information that technically exists online. Volume and quality are not the same thing, and that distinction is now shaping decisions at the highest levels of finance and technology.

Elon Musk: Synthetic data as the only path forward

Among the more provocative positions is that of Elon Musk, who argued in early 2025 that AI firms have exhausted the cumulative sum of human knowledge for training purposes. According to The Guardian (2025), Musk claimed synthetic data is now the only viable way to address the shortage for new models. Researchers remain divided on this point. Concerns about model collapse, where AI trained on AI-generated content degrades in quality over successive iterations, mean synthetic data is viewed as a partial solution at best, not a wholesale replacement for human-generated text.

Data scarcity is not a uniform global problem. Regulatory environments, population size, and digital infrastructure mean that some markets feel the pressure of AI running out of data far sooner than others, and the competitive consequences for e-commerce businesses vary significantly by region.

North America and Europe: regulation accelerates the squeeze

In North America and Europe, stricter data governance frameworks are compounding the scarcity problem. GDPR in Europe and increasingly assertive copyright enforcement across both regions have tightened what AI developers can legally use. Platforms such as Reddit and Stack Overflow have restricted unauthorized AI training data access through commercial API terms, and major news organizations have strengthened licensing agreements, shrinking the pool of freely usable, high-quality web text. According to SAP (2024), this tightening of access is pushing enterprises to invest heavily in proprietary data infrastructure, a significant cost advantage for large players over smaller businesses.

Asia-Pacific: a temporary buffer, not a long-term solution

Asia-Pacific markets benefit from larger populations and comparatively less restrictive data policies, creating a short-term advantage in raw data volume. However, high-quality, structured text in many regional languages remains limited, meaning the quantity advantage does not always translate into training quality.

Emerging markets: structural vulnerability

Businesses in emerging markets face a compounding challenge. Limited first-party data infrastructure means greater dependence on public web data. As that public data becomes increasingly restricted or exhausted, these businesses are left more exposed than their counterparts in data-rich economies.

E-commerce implications across all regions

Global marketplaces including Amazon and Alibaba are already beginning to prioritize sellers whose product data is clean, structured, and machine-readable. Across every region, the competitive advantage is shifting. Price alone is no longer sufficient. Regional sellers who invest in data quality now will be better positioned as AI systems become more selective about the inputs they can effectively use.

Frequently asked questions

Will AI really run out of data by 2026?

According to PBS NewsHour (2024), Epoch AI estimates that public human-written text suitable for large language model training could be exhausted as early as 2026, with a broader window stretching to 2032 and a median forecast of 2028. Whether 2026 proves accurate depends on how aggressively labs continue scaling their models.

Is AI running out of data, or is it just hype?

The concern is grounded in measurable research, not speculation. Multiple independent forecasts, including projections from Epoch AI and Anthropic researchers, point to a genuine ceiling on high-quality public text within this decade.

What happens when AI runs out of training data?

Model improvement through simple data scaling slows significantly. Labs must then rely on synthetic data, proprietary datasets, or architectural innovations to continue advancing model performance.

Can synthetic data solve the AI data shortage?

Synthetic data offers a partial solution, but research suggests models trained predominantly on AI-generated content risk compounding errors over successive generations. It is a bridge, not a permanent fix.

How does the AI data shortage affect e-commerce product visibility?

As AI systems become more selective, product listings with structured, high-quality data are prioritized over thin or inconsistent content. Sellers with poor data quality face reduced discoverability across AI-powered search and recommendation surfaces.

What should e-commerce stores do to prepare for AI data scarcity?

Invest in structured, accurate, and machine-readable product data now, before AI systems become even more discriminating. According to Forbes (2026), authentic, well-structured data is increasingly scarce and therefore increasingly valuable.

Based on our work at Pickastor, businesses that audit and optimize their product data today are measurably better positioned as AI ranking systems grow more selective. The Pickastor AI Optimization Platform is designed specifically to help e-commerce teams build the kind of structured, AI-ready product content that performs as data quality becomes the defining competitive variable.

Is your store ready for AI commerce?

Get your free AI Score - no signup required.

Scan your store for free →