Terminal-Bench
tbench.aiTerminal-Bench is an open-source benchmarking suite designed to evaluate the performance of AI agents in terminal and command-line environments. Developed as a collaboration between Stanford and Laude, it provides standardized tasks across system administration, security, and software engineering to quantify agent mastery. The platform offers leaderboards and detailed metrics to help developers measure and improve their AI tools’ capabilities in real-world technical workflows.
LLM mention score The LLM mention score is the total number of mentions of this brand in different LLM chatbots, normalized to the scale from 0 to 100. You can get actual, non-normalized numbers via the LLM Mention API from DataForSEO.
Normalized 0–100 · last 8 weeks
DataForSEO API
Get LLM mention data of any company via DataForSEO API
Get access to the structured data on keyword, brand, and website mentions in LLMs, including metrics like AI search volume, impressions, and mentions count.
How to get LLM mention data →// Fetch Terminal-Bench mentions POST v3/ai_optimization/llm_mentions/search/live [ { "target": [ { "keyword": "Terminal-Bench", "search_scope": ["any"] } ], "platform": "chat_gpt", "order_by" : ["ai_search_volume,desc"] } ]