Where Do AI Companies Really Get Their Data? 7 Key Sources Revealed

Discover how AI companies source training data through web crawling, licensing deals, and synthetic data. Learn what this means for your ecommerce visibility.

Rihards Ručevics16 min read
Where Do AI Companies Really Get Their Data? 7 Key Sources Revealed
Where do AI companies really get their data? 7 key sources revealed

Introduction: understanding AI data sourcing and why it matters to ecommerce

Most ecommerce merchants assume AI companies simply crawl the web and call it done. The reality is far more complex, and understanding it can give your business a genuine competitive edge. AI companies use multiple data acquisition methods, each with different implications for what gets seen, indexed, and ultimately recommended to consumers.

Why data sourcing is no longer just a technical question

At Pickastor, our analysis shows that ecommerce merchants who understand how AI systems acquire data are significantly better positioned to influence what those systems surface about their products. This is not a passive process. The decisions AI companies make about where to collect data directly shape which brands appear in AI-generated recommendations, product comparisons, and shopping guidance.

According to Bridging the Data Provenance Gap Across Text, Speech and Video (2024), the sourcing landscape for AI training data has grown considerably more diverse, spanning text, speech, video, and structured datasets, each governed by different licensing conditions.

The shift from indiscriminate crawling to licensed datasets

The industry is moving away from broad, indiscriminate web crawling toward purposefully licensed, high-quality datasets. Commercial licensing agreements between AI developers and content owners are expanding rapidly, and data provenance is becoming a strategic differentiator rather than an afterthought.

For ecommerce teams, this shift matters practically. Your product data, structured content, and brand signals are increasingly inputs into AI discovery systems. Merchants who treat their data as a strategic asset, rather than a byproduct of operations, are the ones AI systems will learn from and recommend most reliably.

The three primary data sources AI companies rely on today

AI companies draw from a surprisingly narrow set of foundational channels, even as their models grow more capable. Understanding these three sources gives ecommerce teams a clearer picture of how AI systems are trained, what content they prioritize, and where your product data fits into that ecosystem.

Web crawling and Common Crawl

Web crawling remains the backbone of most large language model training. Common Crawl, the nonprofit archive that scrapes and stores publicly accessible web content, has become a near-universal starting point. The scale is significant: Common Crawl contributes roughly 130 trillion tokens to many training datasets, compared to the 510 trillion tokens that represent a more complete picture of what frontier models require.

That gap matters. It tells you that crawling alone is no longer sufficient, and that AI developers are actively sourcing data elsewhere to fill it. It also tells you something important about visibility: content that is not crawlable, not structured, or not publicly indexed is effectively invisible to these systems. If you want a deeper grounding in what makes data useful to AI, What Is AI Data? A Clear Definition with Practical Examples is a practical starting point.

Commercial licensing agreements

As crawling faces growing legal and ethical scrutiny, AI companies are turning to formal licensing deals with publishers, data vendors, and content platforms. According to the Publishers Association (2025), at least 10 major UK publishers have already signed AI licensing agreements, with 8 more expected to follow. This trend signals a structural shift: premium, rights-cleared content is becoming a distinct and growing data channel.

Synthetic and purpose-built datasets

When real-world data is scarce, legally restricted, or inconsistent in quality, AI developers increasingly generate synthetic data or commission purpose-built datasets. These fill specific capability gaps and reduce provenance uncertainty. For ecommerce teams, this is a useful signal: structured, well-labeled product data is precisely the kind of content that purpose-built datasets are designed to replicate. Everything You Need to Know About Data for AI covers how these distinctions shape AI system behavior in practice.

Expert tip 1: recognize that web crawling is no longer unrestricted

Web crawling was once a free-for-all. AI developers could scrape virtually any public website and feed that content into training pipelines with minimal friction. That era is ending fast, and understanding why matters for anyone whose business depends on how AI systems learn about products, brands, and markets.

The scale of crawling restrictions is larger than most people assume

The numbers here are striking. Research into the C4 dataset, one of the most widely used text corpora in AI training, found that 45% of its content faced some form of crawling restriction, while 5% was fully blocked for scraping. According to the Artificial Intelligence Index Report 2025 (2025), restrictions on data use have increased steadily across major training sources, reflecting a broader shift in how content owners are responding to AI development.

This is not a niche legal concern. It is a structural change in how the web operates.

Robots.txt and terms of service are becoming real enforcement tools

For years, robots.txt files were treated as polite suggestions. AI developers largely ignored them. That is changing as litigation, regulation, and public pressure mount. Publishers, news organizations, and platform operators are now actively enforcing terms of service against unauthorized scraping, and courts in multiple jurisdictions are beginning to weigh in.

This trend connects directly to the broader question of whether AI is running out of data as open web access becomes more constrained.

What this means for ecommerce businesses

For merchants and ecommerce teams, this shift creates a genuine strategic opportunity. Your robots.txt file is not just a technical artifact. It is a policy instrument. Configuring it thoughtfully lets you control which crawlers access your product data, your pricing, and your content, giving you more influence over how your information enters AI training pipelines than most businesses currently exercise.

Expert tip 2: understand the composition of major training datasets

Knowing which platforms and sources dominate AI training datasets gives ecommerce businesses a clearer picture of where their content may already be circulating. The composition of these datasets is more concentrated than most merchants realise, and the attribution picture is murkier still.

More than 70% More than 70% of licenses for popular datasets on GitHub and Hugging Face were unspecified. Nature Machine Intelligence (2024)
608 languages; 798 sources; 659 organizations; 67 countries The analyzed public datasets covered 608 languages, 798 sources, 659 organizations, and 67 countries. Bridging the Data Provenance Gap Across Text, Speech and Video (2024)
Nearly 4,000 datasets The research dataset analyzed nearly 4,000 public datasets spanning text, speech, and video. Bridging the Data Provenance Gap Across Text, Speech and Video (2024)

The platforms that dominate training data

According to Bridging the Data Provenance Gap Across Text, Speech and Video (2024), an audit spanning nearly 4,000 datasets across 608 languages and 798 sources reveals a striking concentration of content origins:

  • Wikipedia accounts for 14.9% of widely used dataset sources, making it the single largest identifiable contributor
  • Undisclosed webpage crawls represent 7.0%, meaning a significant share of training content has no clear origin label
  • Reddit contributes 6.2% and Twitter contributes 4.0%, reflecting how much AI systems have learned from social and community-generated content

For ecommerce teams, this matters because product descriptions, review content, and forum discussions about your brand or category almost certainly fed into these pipelines. Understanding that reality is the first step toward managing it strategically. If you are curious about how AI systems are reshaping data-driven roles more broadly, The Hidden Truth: Will AI Really Take Over Data Science? offers useful context.

The attribution gap that complicates everything

Perhaps the most consequential finding from this research is what remains unknown. Over 70% of dataset licenses are unspecified, meaning the vast majority of training data carries no clear terms governing its use, attribution, or commercial application.

For merchants and agencies, this creates genuine uncertainty. Your product content, blog posts, and category pages may have contributed to AI model training without any record of that contribution existing. This is not a hypothetical concern. It is a structural feature of how most major datasets were assembled, and it has direct implications for how you think about content ownership and visibility in AI-generated outputs.

Expert tip 3: leverage commercial licensing as a data quality differentiator

Commercial licensing is rapidly becoming the clearest signal of data quality in AI development. Where scraped datasets carry ambiguous provenance and contested legal status, licensed content comes with documented origins, defined usage rights, and traceable attribution. For anyone trying to understand where do AI companies get their data, licensing deals represent the most transparent part of the answer.

A formal signing table with two parties exchanging documents, one branded as a publishing house and the other as a technology company, with stacked folders of content agreements visible in the background

The licensing landscape is expanding fast

According to the Publishers Association (2026), 10 major publishers have already signed AI licensing agreements, with at least 8 more expected to follow by the end of 2026. This is not a niche trend. It reflects a structural shift in how AI developers are sourcing content, particularly as legal scrutiny around training data intensifies and regulators push for greater accountability.

The publishers entering these agreements include news organisations, academic publishers, and specialist content providers. Each deal typically establishes clear terms around what content can be used, for which model types, and how attribution is handled. That level of specificity is something scraped datasets simply cannot offer.

Why licensed data matters for ecommerce AI applications

AI companies increasingly prefer licensed data for high-stakes applications, and ecommerce recommendations sit firmly in that category. When a model is generating product suggestions, pricing guidance, or search rankings, the quality and provenance of its training data directly affects output reliability. Licensed content, with its clear chain of custody, reduces the risk of models learning from low-quality, duplicated, or legally contested sources.

For merchants and agencies, this has a practical implication. If your product content is well-structured, authoritative, and consistently maintained, it becomes the kind of material that licensing agreements are built around. Understanding how data analysts and AI specialists are adapting to these shifts can help your team stay ahead of how content quality is evaluated in AI-driven commerce environments.

The shift toward licensing also reinforces a core principle: data quality is not just a technical concern. It is a competitive one.

Common mistakes to avoid when thinking about AI data sourcing

Even experienced teams carry outdated assumptions about how AI systems acquire their training data. These misconceptions can lead to missed opportunities and misplaced priorities, particularly for ecommerce businesses trying to improve their visibility in AI-driven discovery environments.

Assuming web scraping is the whole story

Web scraping is visible and frequently discussed, so it tends to dominate the conversation. But it represents only one layer of a much more complex data ecosystem. According to Bridging the Data Provenance Gap Across Text, Speech and Video (2024), modern AI training pipelines draw from a wide range of source types, including licensed corpora, curated datasets, and purpose-built collections. Treating scraping as the default explanation means overlooking the structured, high-quality inputs that increasingly shape model behaviour.

Ignoring synthetic and purpose-built datasets

Synthetic data is no longer a workaround. It is becoming a primary input for training and fine-tuning AI models, particularly where real-world data is scarce or inconsistently labelled. Teams that dismiss synthetic data as a lesser alternative miss how central it has become to AI development pipelines. If you want to understand who shapes these datasets at a technical level, exploring top AI data labeling companies worth considering this year provides useful context.

Overlooking structured product data and schema markup

In our experience at Pickastor, one of the most consistent gaps we see among ecommerce merchants is neglecting how AI systems actually read and interpret product information. Unstructured listings, missing schema markup, and poorly maintained product feeds are effectively invisible to many AI discovery tools. Optimising for AI readability is not optional if product visibility in AI-generated recommendations matters to your business.

Believing robots.txt provides complete protection

Robots.txt signals intent, but it does not enforce compliance. Many AI data collection processes operate outside the boundaries that robots.txt was designed to govern. Relying on it as a primary control mechanism gives a false sense of security and distracts from more effective strategies.

How ecommerce merchants can optimize for AI data collection

Ecommerce merchants have more control over AI data collection than most realize. By making deliberate technical choices about how product data is structured and presented, you can significantly improve your store's visibility in AI-generated recommendations and ensure the information AI systems discover is accurate and complete.

A developer's screen showing JSON-LD structured data markup overlaid on a product page with pricing and availability fields highlighted

Implement schema.org JSON-LD markup for product data

Structured data is one of the clearest signals you can send to AI crawlers about what your product pages contain. Yet research suggests only 0.77% of pages currently use JSON-LD markup specifically for product data, meaning the vast majority of merchants are invisible to AI systems that rely on structured signals. Adding schema.org Product markup with fields for price, availability, SKU, and brand takes relatively little development effort and delivers outsized returns in AI discoverability. According to Deep Dive: Building an AI-Native Company (2024), the web is actively being restructured around machine-readable formats, and merchants who adapt early capture a compounding advantage.

Ensure product pages are fully crawlable

A product page that loads behind JavaScript rendering, requires login, or blocks key content from crawlers will simply not be indexed by AI systems. Audit your pages to confirm that pricing, availability, identifiers, and descriptions are present in the raw HTML, not injected dynamically after page load.

Create complete, accurate product feeds

AI discovery tools frequently ingest merchant data through product feeds. Treat your feed as a primary channel, not an afterthought. Incomplete attributes, outdated pricing, and missing identifiers all reduce the likelihood that your products appear in AI-generated results.

Use robots.txt strategically, not defensively

Rather than relying on robots.txt as a barrier, use it to selectively allow beneficial AI crawlers while restricting those that offer no commercial value. Understanding how human data teams at companies like OpenAI curate and prioritize sources can help you make more informed decisions about which crawlers to welcome.

Monitor what AI systems actually see

Tools like Pickastor's AI Score give merchants a concrete view of how AI systems interpret their store, identifying gaps between what your pages contain and what AI crawlers actually process. Regular monitoring turns AI data collection from a passive process into one you can actively manage and improve.

Tools and resources for understanding and managing AI data visibility

Understanding where you stand in AI data ecosystems requires the right diagnostic toolkit. Several platforms and free resources give ecommerce merchants concrete visibility into how AI systems discover, interpret, and index their content, turning an otherwise opaque process into something measurable and actionable.

AI Score: see your store through an AI lens

Pickastor's AI Score diagnostic tool reveals exactly what ChatGPT, Google AI Mode, and Perplexity see when they encounter your store. Rather than guessing whether your product pages are being understood correctly, you get a structured breakdown of gaps, strengths, and specific opportunities for improvement.

Pickastor AI Optimization Platform: automate the technical groundwork

The Pickastor AI Optimization Platform handles schema markup injection and product feed generation automatically, removing the technical burden from merchants who lack developer resources. For SMBs especially, this kind of automation is the difference between being visible to AI systems and being invisible.

Free tools worth bookmarking

  • Common Crawl index search: Check whether your content has been captured in one of the most widely used AI training datasets
  • Google Search Console: Identify which pages are crawlable, indexed, and eligible for AI-powered search features
  • Robots.txt analyzers: Audit your crawl directives to confirm you are welcoming the right bots and blocking none unintentionally

For deeper context on how training data shapes AI outputs, the Where Does OpenAI Get Its Data? guide covers the provenance question in detail. As research suggests, ecommerce data is increasingly functioning as an AI-discovery input, making these tools less optional and more foundational to any serious visibility strategy.

Conclusion: taking action to control your ecommerce presence in AI systems

The landscape of AI data sourcing is shifting fast, and merchants who understand that shift will hold a genuine competitive edge. Structured product data, accessible feeds, and clear schema markup have moved from nice-to-have to foundational requirements for any store that wants to appear in AI-driven discovery.

Transparency and provenance are reshaping the rules

According to the Bridging the Data Provenance Gap (2024) research, tracing exactly where AI training data originates is becoming a strategic priority across the industry. For ecommerce teams, this means the quality and clarity of your product data matters not just for search engines but for the AI systems increasingly mediating purchase decisions.

Your next step: audit before you optimize

Start by understanding what AI systems currently see when they encounter your store. Run a structured audit of your feeds, schema markup, and crawl accessibility. Then optimize systematically, addressing gaps in product descriptions, pricing signals, and category structure.

Tools like the Pickastor AI Optimization Platform and its AI Score give you a measurable baseline to work from. For a broader view of the hidden data challenges generative AI faces in 2026, that context will sharpen your optimization priorities considerably.

The merchants who act now will be far better positioned as AI-driven commerce continues to accelerate.

Frequently asked questions

Where do AI companies get their training data?

AI companies draw from a wide range of sources. According to the OECD (2025), AI developers obtain data through commercial licensing agreements, open-data initiatives, web scraping, and repositories such as Common Crawl. Public datasets, social platforms, and licensed content all contribute to the mix.

Do AI companies scrape data from the internet?

Yes, web scraping is a common and significant practice. Sites like Wikipedia, Reddit, and Twitter appear frequently in training datasets, and large-scale crawls of public web pages remain a primary collection method for many foundation models.

What is Common Crawl and how is it used by AI companies?

Common Crawl is a nonprofit repository of petabyte-scale web snapshots made freely available for research and commercial use. According to Common Corpus researchers (2025), large language model training has historically relied heavily on it, though newer datasets increasingly emphasize permissively licensed material.

Do AI companies use copyrighted content to train models?

This is an actively contested legal area. According to Nature Machine Intelligence (2024), more than 70% of licenses for popular datasets were unspecified, meaning the copyright status of much training data remains unclear.

Can website owners stop AI companies from scraping their data?

Website owners can use robots.txt directives and terms-of-service restrictions to signal that scraping is not permitted. Compliance, however, varies considerably across AI developers.

How can ecommerce websites make their product data readable to AI?

Structured data markup, clear product descriptions, and well-organized category hierarchies all improve how AI systems interpret and surface your products. The Pickastor AI Optimization Platform is built specifically to help ecommerce teams audit and improve these signals through its AI Score, giving you a concrete starting point for optimization.

Is your store ready for AI commerce?

Get your free AI Score - no signup required.

Scan your store for free →