# Tasksets Harbor tasksets indexed here. `runnable` tasks have passed their own reference solution on the runner and can be started from the site. | taskset | domain | tasks | runnable | owner | grading | |---|---|---|---|---|---| | [aider_polyglot](https://harnessreport.com/tasks/aider_polyglot.md) | swe | 225 | 6 | | | | [bigcodebench_hard_complete](https://harnessreport.com/tasks/bigcodebench_hard_complete.md) | swe | 145 | 3 | | | | [bird-bench](https://harnessreport.com/tasks/bird-bench.md) | data-sql | 1534 | 3 | | | | [crustbench](https://harnessreport.com/tasks/crustbench.md) | swe | 100 | 3 | | | | [dabstep](https://harnessreport.com/tasks/dabstep.md) | data-sql | 450 | 3 | Adyen | exact-match | | [humanevalfix](https://harnessreport.com/tasks/humanevalfix.md) | swe | 164 | 3 | | | | [quixbugs](https://harnessreport.com/tasks/quixbugs.md) | swe | 80 | 3 | | | | [spider2-dbt](https://harnessreport.com/tasks/spider2-dbt.md) | data-sql | 64 | 3 | | | | [swebench-verified](https://harnessreport.com/tasks/swebench-verified.md) | swe | 500 | 3 | OpenAI & SWE-bench team (Princeton/Stanford) | tests | | [usaco](https://harnessreport.com/tasks/usaco.md) | swe | 304 | 3 | | | | [evoeval](https://harnessreport.com/tasks/evoeval.md) | swe | 100 | 2 | | | | [algotune](https://harnessreport.com/tasks/algotune.md) | swe | 154 | 1 | | | | [aa-lcr](https://harnessreport.com/tasks/aa-lcr.md) | reasoning-knowledge | 99 | 0 | Artificial Analysis | llm-judge | | [abc-bench](https://harnessreport.com/tasks/abc-bench.md) | life-sciences | 224 | 0 | | mixed | | [ace-bench](https://harnessreport.com/tasks/ace-bench.md) | enterprise-ops | 973 | 0 | Mercor | rubric | | [acebench-normal](https://harnessreport.com/tasks/acebench-normal.md) | enterprise-ops | 823 | 0 | Mercor | rubric | | [acebench-special](https://harnessreport.com/tasks/acebench-special.md) | enterprise-ops | 150 | 0 | Mercor | rubric | | [ade-bench](https://harnessreport.com/tasks/ade-bench.md) | data-sql | 48 | 0 | | | | [aime](https://harnessreport.com/tasks/aime.md) | science-math | 60 | 0 | Mathematical Association of America | exact-match | | [arc_agi_1](https://harnessreport.com/tasks/arc_agi_1.md) | reasoning-knowledge | 400 | 0 | ARC Prize Foundation | exact-match | | [arc_agi_2](https://harnessreport.com/tasks/arc_agi_2.md) | reasoning-knowledge | 167 | 0 | ARC Prize Foundation | exact-match | | [autocodebench](https://harnessreport.com/tasks/autocodebench.md) | swe | 200 | 0 | | | | [bfcl](https://harnessreport.com/tasks/bfcl.md) | tool-use | 3641 | 0 | UC Berkeley (Gorilla) | state-check | | [bfcl_parity](https://harnessreport.com/tasks/bfcl_parity.md) | tool-use | 123 | 0 | UC Berkeley (Gorilla) | state-check | | [bigcodebench_hard_instruct](https://harnessreport.com/tasks/bigcodebench_hard_instruct.md) | swe | 145 | 0 | | | | [bixbench](https://harnessreport.com/tasks/bixbench.md) | life-sciences | 205 | 0 | FutureHouse | mixed | | [bixbench-cli](https://harnessreport.com/tasks/bixbench-cli.md) | life-sciences | 205 | 0 | FutureHouse | mixed | | [clbench](https://harnessreport.com/tasks/clbench.md) | reasoning-knowledge | 1899 | 0 | Tencent Hunyuan | rubric | | [codepde](https://harnessreport.com/tasks/codepde.md) | swe | 5 | 0 | | | | [compilebench](https://harnessreport.com/tasks/compilebench.md) | swe | 15 | 0 | | | | [cooperbench](https://harnessreport.com/tasks/cooperbench.md) | reasoning-knowledge | 652 | 0 | | | | [crmarena](https://harnessreport.com/tasks/crmarena.md) | customer-service | 1170 | 0 | Salesforce AI Research | exact-match | | [cybergym](https://harnessreport.com/tasks/cybergym.md) | security | 6028 | 0 | UC Berkeley (Sunblaze / RDI) | state-check | | [dacode](https://harnessreport.com/tasks/dacode.md) | data-sql | 479 | 0 | | | | [deepsynth](https://harnessreport.com/tasks/deepsynth.md) | web-research | 40 | 0 | | | | [deveval](https://harnessreport.com/tasks/deveval.md) | swe | 63 | 0 | | | | [devopsgym](https://harnessreport.com/tasks/devopsgym.md) | devops-sre | 733 | 0 | | | | [ds1000](https://harnessreport.com/tasks/ds1000.md) | data-sql | 1000 | 0 | XLANG Lab (HKU) | tests | | [featbench](https://harnessreport.com/tasks/featbench.md) | swe | 156 | 0 | | | | [featurebench](https://harnessreport.com/tasks/featurebench.md) | swe | 200 | 0 | LiberCoders | tests | | [featurebench-lite](https://harnessreport.com/tasks/featurebench-lite.md) | swe | 30 | 0 | LiberCoders | tests | | [featurebench-lite-modal](https://harnessreport.com/tasks/featurebench-lite-modal.md) | swe | 30 | 0 | LiberCoders | tests | | [featurebench-modal](https://harnessreport.com/tasks/featurebench-modal.md) | swe | 200 | 0 | LiberCoders | tests | | [financeagent](https://harnessreport.com/tasks/financeagent.md) | finance | 50 | 0 | Vals AI | llm-judge | | [financeagent_terminal](https://harnessreport.com/tasks/financeagent_terminal.md) | finance | 50 | 0 | Vals AI | llm-judge | | [frontier-cs](https://harnessreport.com/tasks/frontier-cs.md) | swe | 172 | 0 | | tests | | [gaia](https://harnessreport.com/tasks/gaia.md) | tool-use | 165 | 0 | Meta + Hugging Face | exact-match | | [gaia2](https://harnessreport.com/tasks/gaia2.md) | customer-service | 800 | 0 | Meta | state-check | | [gaia2-cli](https://harnessreport.com/tasks/gaia2-cli.md) | customer-service | 800 | 0 | Meta | state-check | | [gdb](https://harnessreport.com/tasks/gdb.md) | office-docs | 78 | 0 | | | | [gdb-hub](https://harnessreport.com/tasks/gdb-hub.md) | office-docs | 33786 | 0 | | | | [gpqa-diamond](https://harnessreport.com/tasks/gpqa-diamond.md) | science-math | 198 | 0 | NYU / Cohere / Anthropic (Rein et al.) | exact-match | | [gso](https://harnessreport.com/tasks/gso.md) | swe | 102 | 0 | | | | [hle](https://harnessreport.com/tasks/hle.md) | reasoning-knowledge | 2500 | 0 | Center for AI Safety & Scale AI | llm-judge | | [ineqmath](https://harnessreport.com/tasks/ineqmath.md) | science-math | 100 | 0 | | | | [kramabench](https://harnessreport.com/tasks/kramabench.md) | data-sql | 104 | 0 | | | | [kumo](https://harnessreport.com/tasks/kumo.md) | reasoning-knowledge | 5300 | 0 | | | | [labbench](https://harnessreport.com/tasks/labbench.md) | life-sciences | 181 | 0 | FutureHouse | exact-match | | [lawbench](https://harnessreport.com/tasks/lawbench.md) | legal | 1000 | 0 | | | | [livecodebench](https://harnessreport.com/tasks/livecodebench.md) | swe | 100 | 0 | LiveCodeBench team (UC Berkeley/MIT/Cornell) | tests | | [llmsr-bench](https://harnessreport.com/tasks/llmsr-bench.md) | science-math | 240 | 0 | | | | [locomo](https://harnessreport.com/tasks/locomo.md) | reasoning-knowledge | 10 | 0 | | | | [medagentbench](https://harnessreport.com/tasks/medagentbench.md) | healthcare | 300 | 0 | Stanford | state-check | | [ml_dev_bench](https://harnessreport.com/tasks/ml_dev_bench.md) | ml-research | 33 | 0 | | | | [mlgym-bench](https://harnessreport.com/tasks/mlgym-bench.md) | ml-research | 11 | 0 | Meta | tests | | [mlgym-bench-hub](https://harnessreport.com/tasks/mlgym-bench-hub.md) | ml-research | 12 | 0 | Meta | tests | | [mmau](https://harnessreport.com/tasks/mmau.md) | multimodal | 1000 | 0 | | | | [mmmlu](https://harnessreport.com/tasks/mmmlu.md) | science-math | 150 | 0 | | | | [multi-swe-bench](https://harnessreport.com/tasks/multi-swe-bench.md) | swe | 1601 | 0 | ByteDance Seed | tests | | [omnimath](https://harnessreport.com/tasks/omnimath.md) | science-math | 4428 | 0 | | | | [osworld-external-credentials](https://harnessreport.com/tasks/osworld-external-credentials.md) | other | 8 | 0 | | | | [osworld-verified](https://harnessreport.com/tasks/osworld-verified.md) | computer-use | 361 | 0 | XLANG Lab (HKU) | state-check | | [pixiu](https://harnessreport.com/tasks/pixiu.md) | finance | 435 | 0 | | | | [programbench](https://harnessreport.com/tasks/programbench.md) | swe | 200 | 0 | Meta FAIR (facebookresearch) | tests | | [qcircuitbench](https://harnessreport.com/tasks/qcircuitbench.md) | science-math | 28 | 0 | | | | [reasoning-gym](https://harnessreport.com/tasks/reasoning-gym.md) | reasoning-knowledge | 576 | 0 | | | | [refav](https://harnessreport.com/tasks/refav.md) | engineering-industrial | 1500 | 0 | | | | [replicationbench](https://harnessreport.com/tasks/replicationbench.md) | web-research | 90 | 0 | | | | [research-code-bench](https://harnessreport.com/tasks/research-code-bench.md) | swe | 212 | 0 | | | | [rexbench](https://harnessreport.com/tasks/rexbench.md) | ml-research | 2 | 0 | | | | [satbench](https://harnessreport.com/tasks/satbench.md) | reasoning-knowledge | 2100 | 0 | | | | [scicode](https://harnessreport.com/tasks/scicode.md) | science-math | 80 | 0 | SciCode team (UIUC et al.) | tests | | [scienceagentbench](https://harnessreport.com/tasks/scienceagentbench.md) | science-math | 102 | 0 | OSU NLP | tests | | [seal0](https://harnessreport.com/tasks/seal0.md) | reasoning-knowledge | 111 | 0 | | | | [simpleqa](https://harnessreport.com/tasks/simpleqa.md) | reasoning-knowledge | 4326 | 0 | OpenAI | mixed | | [sldbench](https://harnessreport.com/tasks/sldbench.md) | science-math | 8 | 0 | | | | [spreadsheetbench-verified](https://harnessreport.com/tasks/spreadsheetbench-verified.md) | office-docs | 400 | 0 | | | | [strongreject](https://harnessreport.com/tasks/strongreject.md) | ai-safety | 150 | 0 | | | | [swe-lancer](https://harnessreport.com/tasks/swe-lancer.md) | swe | 463 | 0 | OpenAI | tests | | [swebench_multilingual](https://harnessreport.com/tasks/swebench_multilingual.md) | swe | 300 | 0 | SWE-bench team (Princeton/Stanford) | tests | | [swebenchpro](https://harnessreport.com/tasks/swebenchpro.md) | swe | 731 | 0 | Scale AI | tests | | [swegym](https://harnessreport.com/tasks/swegym.md) | swe | 2438 | 0 | | | | [swegym-lite](https://harnessreport.com/tasks/swegym-lite.md) | swe | 230 | 0 | | | | [swesmith](https://harnessreport.com/tasks/swesmith.md) | swe | 100 | 0 | | | | [swtbench-verified](https://harnessreport.com/tasks/swtbench-verified.md) | swe | 433 | 0 | | | | [tau3-bench](https://harnessreport.com/tasks/tau3-bench.md) | customer-service | 375 | 0 | Sierra | state-check | | [textarena](https://harnessreport.com/tasks/textarena.md) | games | 62 | 0 | | | | [theagentcompany](https://harnessreport.com/tasks/theagentcompany.md) | enterprise-ops | 174 | 0 | | | | [webgen-bench](https://harnessreport.com/tasks/webgen-bench.md) | web-research | 101 | 0 | | | | [widesearch](https://harnessreport.com/tasks/widesearch.md) | web-research | 200 | 0 | ByteDance Seed | mixed | --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp