SWE-bench
π» Coding agents Real-world GitHub issues from popular Python repositories. The agent must edit code so the repository's own test suite passes.
Pros: Highly realistic; directly measures production software engineering ability.
Cons: Expensive to run; requires sandboxed execution and repo setup; not all issues are agent-friendly.
Best fit: Code-generation agents, AI software engineers, and autonomous PR bots.
Effort: High Cost: $500 β $5,000 per evaluation run
View benchmark β
HumanEval
π» Coding agents 164 hand-written Python function-completion problems with unit-test validation.
Pros: Cheap, fast, and widely cited; great for comparing pass@1 coding skill across models.
Cons: Short functions only; may be memorized by large models; does not test repo-level work.
Best fit: Quick model comparisons for code completion and small-function generation.
Effort: Low Cost: $5 β $50 per run
View benchmark β
MBPP (Mostly Basic Python Problems)
π» Coding agents Crowd-sourced Python programming problems with test cases, from beginner to intermediate difficulty.
Pros: Larger and more varied than HumanEval; includes docstrings and natural-language prompts.
Cons: Still focused on small, self-contained functions; less realistic than repo-level benchmarks.
Best fit: Entry-level coding assistants and educational coding agents.
Effort: Low Cost: $10 β $100 per run
View benchmark β
GAIA
π§ Reasoning & knowledge Multi-step question answering that requires reasoning, web browsing, file parsing, and tool use to answer real-world questions.
Pros: Very hard to game; rewards agents that can plan, search, and synthesize evidence.
Cons: Slow and expensive; requires web access; answers can be brittle to search-result changes.
Best fit: General-purpose assistant agents, research agents, and agents with tool access.
Effort: High Cost: $300 β $2,000 per run
View benchmark β
MMLU
π§ Reasoning & knowledge 57 multiple-choice tasks covering STEM, humanities, social sciences, and professional knowledge.
Pros: Standardized, cheap, and widely reported; good baseline for world knowledge.
Cons: Static multiple-choice benchmark; does not measure multi-step reasoning or agentic behavior.
Best fit: Foundation-model selection and knowledge-heavy assistant agents.
Effort: Low Cost: $10 β $100 per run
View benchmark β
LiveBench
π§ Reasoning & knowledge Continuously updated reasoning tasks drawn from recent data sources, designed to avoid contamination.
Pros: Harder to memorize; stays current; covers math, reasoning, language, and instruction following.
Cons: Less established than MMLU or HumanEval; questions can be noisy.
Best fit: Agents that must stay current or claim strong general reasoning.
Effort: Medium Cost: $50 β $500 per run
View benchmark β
WebArena
π Web agents Independent, realistic web tasks across e-commerce, social forums, software development, maps, and admin panels.
Pros: Simulates real websites; measures end-to-end task completion in a live environment.
Cons: Requires hosted environments; can be flaky due to website changes; expensive to run at scale.
Best fit: Browser agents, web assistants, and autonomous research agents.
Effort: High Cost: $500 β $3,000 per run
View benchmark β
WebVoyager
π Web agents Visual web navigation tasks where the agent must interact with live websites through screenshots and predicted actions.
Pros: Tests vision + web interaction; closer to how humans browse visually.
Cons: Latency is high; requires screenshot rendering; brittle to UI redesigns.
Best fit: Vision-enabled browser agents and multimodal assistants.
Effort: High Cost: $400 β $2,500 per run
View benchmark β
Mind2Web
π Web agents Diverse web tasks on real websites across multiple domains, annotated with expert demonstrations.
Pros: Large coverage of web domains; includes both action prediction and task completion metrics.
Cons: Benchmark is static; newer website versions may diverge from the dataset.
Best fit: Generalist web agents and UI automation tools.
Effort: Medium Cost: $100 β $800 per run
View benchmark β
OSWorld
π₯οΈ Desktop / OS agents Open-ended tasks inside a real Ubuntu desktop: file management, browser use, office apps, and system configuration.
Pros: Real OS environment; captures messy, long-horizon computer use.
Cons: Very expensive and slow; needs VM infrastructure; high variance across runs.
Best fit: Desktop automation agents, personal assistants, and OS-control agents.
Effort: Very High Cost: $1,000 β $10,000 per run
View benchmark β
Windows Agent Arena
π₯οΈ Desktop / OS agents Tasks on a live Windows desktop spanning settings, apps, and cross-application workflows.
Pros: Tests the most common consumer OS; surfaces enterprise automation value.
Cons: Windows-specific licensing and VM setup; evaluation pipeline is complex.
Best fit: Enterprise desktop-automation agents and Windows-centric copilots.
Effort: Very High Cost: $1,500 β $8,000 per run
View benchmark β
AgentBench
π€ Multi-agent systems Eight environments (OS, database, knowledge graph, digital card game, web shopping, household, etc.) that exercise decision making, tool use, and multi-turn interaction.
Pros: Broad multi-environment coverage; compares agents across many dimensions.
Cons: Some environments are synthetic; setup can be heavy; scoring is nuanced.
Best fit: Generalist agents, tool-use agents, and multi-environment systems.
Effort: High Cost: $300 β $2,500 per run
View benchmark β
BenchMAX
π€ Multi-agent systems Multi-agent collaboration tasks such as negotiation, team problem solving, and resource allocation.
Pros: Targets the unique failure modes of multi-agent communication and coordination.
Cons: Smaller community; less standardized than single-agent benchmarks.
Best fit: Multi-agent orchestrators, agent teams, and collaborative AI systems.
Effort: Medium Cost: $200 β $1,500 per run
View benchmark β
GSM8K
π― Domain-specific agents Grade-school math word problems that require multi-step arithmetic reasoning.
Pros: Clean, cheap, and widely used; strong signal for chain-of-thought reasoning.
Cons: Narrow domain; not representative of real agent deployments.
Best fit: Tutoring agents, math copilots, and reasoning-heavy assistants.
Effort: Low Cost: $5 β $50 per run
View benchmark β
CYBER / CTF Evals
π― Domain-specific agents Capture-the-flag and cybersecurity challenges ranging from reconnaissance to exploit development.
Pros: Directly measures offensive-security and vulnerability-research capability.
Cons: High stakes; requires isolated sandbox; not suitable for public leaderboards without review.
Best fit: Security research agents, red-team tools, and vulnerability-assessment copilots.
Effort: Very High Cost: $1,000 β $10,000+ per run
View benchmark β
SciWorld
π― Domain-specific agents Open-ended science experiments in a text-based simulation where agents must design and execute experiments.
Pros: Tests scientific reasoning and long-horizon planning in a controlled simulator.
Cons: Niche; requires domain expertise to interpret results; simulator limitations affect generalization.
Best fit: Scientific research assistants, lab-automation agents, and hypothesis-testing systems.
Effort: High Cost: $500 β $4,000 per run
View benchmark β No benchmarks match those filters. Try widening your effort or budget.