SWE-bench

swebench.com

SWE-bench is a benchmarking framework designed to evaluate the software engineering capabilities of large language models. It provides standardized leaderboards and verified datasets to measure how effectively AI agents can resolve real-world GitHub issues. The platform supports comparison across various open-source and proprietary models using consistent evaluation harnesses.

LLM mention score The LLM mention score is the total number of mentions of this brand in different LLM chatbots, normalized to the scale from 0 to 100. You can get actual, non-normalized numbers via the LLM Mention API from DataForSEO.

Normalized 0–100 · last 8 weeks

DataForSEO API

Get LLM mention data of any company via DataForSEO API

Get access to the structured data on keyword, brand, and website mentions in LLMs, including metrics like AI search volume, impressions, and mentions count. 

How to get LLM mention data →
// Fetch SWE-bench mentions
POST v3/ai_optimization/llm_mentions/search/live
[
    {
        "target": [
            {
                "keyword": "SWE-bench",
                "search_scope": ["any"]
            }
        ],
        "platform": "chat_gpt",
        "order_by" : ["ai_search_volume,desc"]
    }
]