Benchmarks
15 evaluation frameworks and benchmark suites
15 evaluation frameworks and benchmark suites
License
Language
Type
| Name↑ | Section | Lang | License | Type | Description |
|---|---|---|---|---|---|
| AgentBench (THUDM) | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | Apache-2.0 | BENCHMARK | Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licensed |
| GAIA | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Real-world multi-step agentic benchmark |
| HELM (Stanford) | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Holistic Evaluation of Language Models |
| LiveBench | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Contamination-free LLM benchmark with objective ground-truth scoring. ICLR 2025 spotlight paper featuring frequently-updated questions from recent sources. Tests math, coding, reasoning, language, instruction following, and data analysis |
| lm-evaluation-harness (EleutherAI) | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | De-facto standard for generative model evaluation |
| MLE-bench (OpenAI) | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | MIT | BENCHMARK | Benchmark for measuring how well AI agents perform at machine learning engineering. Evaluates agents on 75 Kaggle competitions covering diverse ML tasks. MIT licensed |
| MLPerf Inference | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Industry-standard ML inference benchmarks with reference implementations for AI accelerators |
| MLPerf Training | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | Apache-2.0 | BENCHMARK | Industry-standard ML training benchmarks from MLCommons. Reference implementations for training AI models at scale across image classification, object detection, NLP, and recommendation tasks. Apache 2.0 licensed |
| MMLU-Pro / GPQA | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | More challenging MMLU-style benchmark suite for evaluating advanced language models with expert-level reasoning questions |
| OpenCompass | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Evaluation platform for benchmarking language and multimodal models across large benchmark suites |
| PinchBench | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | MIT | BENCHMARK | Benchmarking system for evaluating LLM models as OpenClaw coding agents. Built with Rust by the kilo.ai team. MIT licensed |
| SWE-bench | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Evaluates LLMs on real-world GitHub issues from 15+ Python repositories |
| SWE-rebench (Nebius) | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | — | BENCHMARK | Continuously updated benchmark with 21,000+ real-world SWE tasks for evaluating agentic LLMs. Decontaminated, mined from GitHub |
| Vectara Hallucination Leaderboard | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | Apache-2.0 | BENCHMARK | Leaderboard comparing LLM performance at producing hallucinations when summarizing short documents. Systematic evaluation of factual consistency across major models. Apache 2.0 licensed |
| VLMEvalKit | Evaluation, Benchmarks & DatasetsBenchmark Suites | — | Apache-2.0 | BENCHMARK | Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed |