# mlgym-bench - domain: ml-research - n_tasks: 11 - n_runnable: 0 - owner_org: Meta - grading: tests - task_kind: open-ended ML research tasks in a Gym env (CV, NLP, RL, game theory) - url_repo: https://github.com/facebookresearch/MLGym - url_paper: https://arxiv.org/abs/2502.14499 ## Tasks | task | difficulty | language | runnable | harnesses tried | |---|---|---|---|---| | [mlgym-battle-of-sexes](https://harnessreport.com/tasks/mlgym-bench/mlgym-battle-of-sexes.md) | medium | | | 0 | | [mlgym-blotto](https://harnessreport.com/tasks/mlgym-bench/mlgym-blotto.md) | medium | | | 0 | | [mlgym-image-classification-cifar10](https://harnessreport.com/tasks/mlgym-bench/mlgym-image-classification-cifar10.md) | medium | | | 0 | | [mlgym-image-classification-cifar10-l1](https://harnessreport.com/tasks/mlgym-bench/mlgym-image-classification-cifar10-l1.md) | medium | | | 0 | | [mlgym-image-classification-f-mnist](https://harnessreport.com/tasks/mlgym-bench/mlgym-image-classification-f-mnist.md) | medium | | | 0 | | [mlgym-prisoners-dilemma](https://harnessreport.com/tasks/mlgym-bench/mlgym-prisoners-dilemma.md) | medium | | | 0 | | [mlgym-regression-kaggle-house-price](https://harnessreport.com/tasks/mlgym-bench/mlgym-regression-kaggle-house-price.md) | medium | | | 0 | | [mlgym-regression-kaggle-house-price-l1](https://harnessreport.com/tasks/mlgym-bench/mlgym-regression-kaggle-house-price-l1.md) | medium | | | 0 | | [mlgym-rl-meta-maze-misc](https://harnessreport.com/tasks/mlgym-bench/mlgym-rl-meta-maze-misc.md) | medium | | | 0 | | [mlgym-rl-mountain-car-continuous](https://harnessreport.com/tasks/mlgym-bench/mlgym-rl-mountain-car-continuous.md) | medium | | | 0 | | [mlgym-titanic](https://harnessreport.com/tasks/mlgym-bench/mlgym-titanic.md) | medium | | | 0 | --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp