SYS NOMINALLISTINGS 1305 ▲SERVERS 948CLIENTS 103AI AGENTS 240SKILLS 05PLUGINS 03RULES 03EVALS 03
All leaderboards

Evals & benchmarks leaderboard.

Benchmarks, test sets and evaluation harnesses that measure what agents and models can do. Browse all of them at /evals.

Ranked by public GitHub stars per day since the repository was created. Repositories under 14 days old or 10 stars are not ranked.

Updated Oct 5, 2026
#ListingCategoryStars / dayStarsSince
1SWE-bench
Benchmark that scores language models on resolving real GitHub issues by producing a working patch.
Developer tools5.46kOct 2023
2AgentBench
A benchmark that evaluates LLMs as agents across eight environments, from operating systems and databases to web shopping and browsing.
AI & language models3.23.8kJul 2023
3τ-bench
Benchmark for tool-agent-user interaction: an agent with API tools and policies serves a simulated user in airline and retail domains.
AI & language models1.71.5kJun 2024