Common Crawl

commoncrawl.org

Common Crawl is a non-profit organization that maintains an open, free repository of web crawl data for public access and analysis. Since 2007, the organization has provided researchers and developers with wholesale extraction capabilities for large-scale datasets spanning hundreds of billions of web pages. Its primary value proposition lies in democratizing access to web data to support academic research, AI model training, and computational linguistics.

LLM mention score The LLM mention score is the total number of mentions of this brand in different LLM chatbots, normalized to the scale from 0 to 100. You can get actual, non-normalized numbers via the LLM Mention API from DataForSEO.

Normalized 0–100 · last 8 weeks

DataForSEO API

Get LLM mention data of any company via DataForSEO API

Get access to the structured data on keyword, brand, and website mentions in LLMs, including metrics like AI search volume, impressions, and mentions count. 

How to get LLM mention data →
// Fetch Common Crawl mentions
POST v3/ai_optimization/llm_mentions/search/live
[
    {
        "target": [
            {
                "keyword": "Common Crawl",
                "search_scope": ["any"]
            }
        ],
        "platform": "chat_gpt",
        "order_by" : ["ai_search_volume,desc"]
    }
]