OpenRuna
Sign in

AgentBench (THUDM)

BENCHMARK

Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licensed

View on GitHub

Overview

AgentBench (THUDM) is a free benchmark on OpenRuna. Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licens

What this benchmark does

Looking for a dependable benchmark? "AgentBench (THUDM)" gives you a tested starting point instead of a blank prompt box. Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licensed It is catalogued next to similar resources on OpenRuna, so the rest of the toolkit you need is close by. Open a new conversation and paste it in, wire it into an agent, or keep it in your team's prompt library.

Use cases

  • Fork it as a baseline and layer in your own project context, constraints, and examples.
  • Keep it in a shared library as the canonical version of this benchmark for your organisation.
  • Use "AgentBench (THUDM)" when you need a repeatable benchmark for professional work without rewriting instructions every time.
  • Hand "AgentBench (THUDM)" to a new teammate so their benchmark output matches your team's quality bar from day one.

Example output

Ask the model to apply "AgentBench (THUDM)" to your scenario and it returns a structured answer — clear sections, actionable steps, and assumptions stated upfront — ready to paste into docs, tickets, or code comments. Add one example of your own and the output quality jumps noticeably.

Tips by platform

Claude

In Claude, paste the full benchmark as your first message or add it to Project instructions, then ask Claude to confirm assumptions before it executes. For longer benchmarks, iterate inside the artifact panel.

ChatGPT

ChatGPT responds well when you paste this benchmark and immediately give one concrete example of your input. Use a reasoning-capable model for multi-step work.

Cursor

In Cursor, lift the key instructions from this benchmark into .cursorrules or a SKILL.md file, then reference it in Agent mode with @ mentions. Keep the title in a comment so teammates can find it on OpenRuna.

Frequently asked questions

What is "AgentBench (THUDM)"?
It is a benchmark listed on OpenRuna — Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licensed You can copy and adapt it for ChatGPT, Claude, Cursor, or any other AI assistant.
Is "AgentBench (THUDM)" free to use?
Most OpenRuna resources are open or CC0-licensed. Check the license shown on this page before commercial use; premium collections are clearly marked as such.
How do I get the best results from this benchmark?
Replace any placeholders, add your project context, and ask the model to confirm its assumptions first. Iterate over 2–3 follow-up turns rather than expecting a perfect first response.
Does "AgentBench (THUDM)" work with both Claude and ChatGPT?
Yes — it is model-agnostic text, so it runs on Claude, ChatGPT, Gemini, and Cursor. The tips on this page cover each of those assistants specifically.
Where can I find resources related to "AgentBench (THUDM)"?
Scroll to the Related resources section on this page, or open the matching category hub on OpenRuna to find connected prompts, tools, agents, and datasets in the same topic area.

Related resources