Evals & benchmarks leaderboard.
Benchmarks, test sets and evaluation harnesses that measure what agents and models can do. Browse all of them at /evals.
| # | Listing | Category | Stars / day | Stars | Since |
|---|---|---|---|---|---|
| 1 | SWE-bench Benchmark that scores language models on resolving real GitHub issues by producing a working patch. | Developer tools | 5.4 | 6k | Oct 2023 |
| 2 | AgentBench A benchmark that evaluates LLMs as agents across eight environments, from operating systems and databases to web shopping and browsing. | AI & language models | 3.2 | 3.8k | Jul 2023 |
| 3 | τ-bench Benchmark for tool-agent-user interaction: an agent with API tools and policies serves a simulated user in airline and retail domains. | AI & language models | 1.7 | 1.5k | Jun 2024 |