MLE-bench (OpenAI)
BENCHMARKBenchmark for measuring how well AI agents perform at machine learning engineering. Evaluates agents on 75 Kaggle competitions covering diverse ML tasks. MIT licensed
OpenRuna's Data & Analytics hub is your starting point for data analytics AI prompts in 2026. Data analysis, business intelligence, statistics, visualisation, and ML workflows. Browse copy-ready prompts, open-source tools, agent templates, datasets, and skills — all organized for developers, builders, and teams who want practical AI resources without sifting through generic lists. Use the filters below to explore prompts, tools, agents, datasets, and workflows in this category. Every resource links into OpenRuna's connected graph so you can discover related assets quickly. Whether you use ChatGPT, Claude, Cursor, or Gemini, these data analytics AI prompts are selected for real-world data & analytics workflows.
34 resources
Benchmark for measuring how well AI agents perform at machine learning engineering. Evaluates agents on 75 Kaggle competitions covering diverse ML tasks. MIT licensed
Benchmarking system for evaluating LLM models as OpenClaw coding agents. Built with Rust by the kilo.ai team. MIT licensed
Comprehensive benchmark to evaluate LLMs as agents across 8 diverse environments including household, web shopping, OS interaction, and database tasks. ICLR 2024. Apache 2.0 licensed
Continuously updated benchmark with 21,000+ real-world SWE tasks for evaluating agentic LLMs. Decontaminated, mined from GitHub
Leaderboard comparing LLM performance at producing hallucinations when summarizing short documents. Systematic evaluation of factual consistency across major models. Apache 2.0 licensed
Open-source evaluation toolkit for large multi-modality models (LMMs). Supports 220+ LMMs and 80+ benchmarks including MMMU, MathVista, and ChartQA. Powers the OpenVLM Leaderboard. Apache 2.0 licensed
Industry-standard ML training benchmarks from MLCommons. Reference implementations for training AI models at scale across image classification, object detection, NLP, and recommendation tasks. Apache 2.0 licensed
Industry-standard ML inference benchmarks with reference implementations for AI accelerators
Evaluation platform for benchmarking language and multimodal models across large benchmark suites
Real-world multi-step agentic benchmark
Evaluates LLMs on real-world GitHub issues from 15+ Python repositories
More challenging MMLU-style benchmark suite for evaluating advanced language models with expert-level reasoning questions
Holistic Evaluation of Language Models
De-facto standard for generative model evaluation
Contamination-free LLM benchmark with objective ground-truth scoring. ICLR 2025 spotlight paper featuring frequently-updated questions from recent sources. Tests math, coding, reasoning, language, instruction following, and data analysis