SYS NOMINALLISTINGS 1305 ▲SERVERS 948CLIENTS 103AI AGENTS 240SKILLS 05PLUGINS 03RULES 03EVALS 03
All leaderboards

Evals & benchmarks leaderboard.

Benchmarks, test sets and evaluation harnesses that measure what agents and models can do. Browse all of them at /evals.

Ranked by public GitHub stars.

Updated Oct 5, 2026
#ListingCategoryStarsForksLast push
1SWE-bench
Benchmark that scores language models on resolving real GitHub issues by producing a working patch.
Developer tools6k1kSep 18, 2026
2AgentBench
A benchmark that evaluates LLMs as agents across eight environments, from operating systems and databases to web shopping and browsing.
AI & language models3.8k287Feb 8, 2026
3τ-bench
Benchmark for tool-agent-user interaction: an agent with API tools and policies serves a simulated user in airline and retail domains.
AI & language models1.5k218Mar 18, 2026