SYS NOMINALLISTINGS 1305 ▲SERVERS 948CLIENTS 103AI AGENTS 240SKILLS 05PLUGINS 03RULES 03EVALS 03
All tools
Ag
EVAL

AgentBench

A benchmark that evaluates LLMs as agents across eight environments, from operating systems and databases to web shopping and browsing.

AI & language modelsevalbenchmarkllm-agentsmulti-environment
Visit project
Claim this listing

About AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR’24). AgentBench is the first benchmark designed to evaluate LLM-as-Agent across a diverse spectrum of different environments. It covers eight environments: operating system, database, knowledge graph, digital card game, lateral thinking puzzles, house-holding, web shopping and web browsing. Paper: arXiv:2308.03688. Details were read from the publisher’s repository and documentation; nothing was installed or executed.

What you can do

  • Eight agent environments
  • LLM-as-Agent evaluation
  • Public leaderboard

Benchmark

Format
Evaluation set or harness
Works with
Not specified by the publisher
License
Apache-2.0

How to use it

How do I run AgentBench?

Follow the publisher’s documentation to get the data and run the harness against your agent or model.

What does AgentBench measure?

Eight agent environments; LLM-as-Agent evaluation; Public leaderboard.

Is it free?

Open source; free to run yourself.

Questions and answers

What is AgentBench?

AgentBench is an eval in the AI & language models category. A benchmark that evaluates LLMs as agents across eight environments, from operating systems and databases to web shopping and browsing.

How do I run AgentBench?

Follow the publisher's documentation at https://arxiv.org/abs/2308.03688 to get the data and run the harness against your agent or model.

Is AgentBench free?

Open source; free to run yourself.

What can AgentBench do?

According to the published details: Eight agent environments; LLM-as-Agent evaluation; Public leaderboard.

Where is AgentBench published?

Its homepage is https://github.com/THUDM/AgentBench and its source repository is https://github.com/THUDM/AgentBench. RUAGENTIC read these details from Publisher repository on October 1, 2026 and did not install or run the project.

Sources and checks

View the source

Source information collected October 1, 2026.

Badge and embed

Show this listing on your website or README. The badge and the card link back to this page.

Listed on RUAGENTICAgentBench on RUAGENTIC
Markdown badge
[![Listed on RUAGENTIC](https://ruagentic.com/badge/agentbench.svg)](https://ruagentic.com/tools/agentbench)
HTML badge
<a href="https://ruagentic.com/tools/agentbench"><img src="https://ruagentic.com/badge/agentbench.svg" alt="Listed on RUAGENTIC" height="28"></a>
Markdown card
[![AgentBench on RUAGENTIC](https://ruagentic.com/embed/agentbench.svg)](https://ruagentic.com/tools/agentbench)
HTML embed
<iframe src="https://ruagentic.com/embed/agentbench" width="480" height="150" style="border:0;max-width:100%" loading="lazy" title="AgentBench on RUAGENTIC"></iframe>