# ml_dev_bench - domain: ml-research - n_tasks: 33 - n_runnable: 0 ## Tasks | task | difficulty | language | runnable | harnesses tried | |---|---|---|---|---| | [ml_dev_bench_basic_vision_finetuning](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_basic_vision_finetuning.md) | hard | | | 0 | | [ml_dev_bench_bert_eval_debug](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_bert_eval_debug.md) | hard | | | 0 | | [ml_dev_bench_boolq_performance](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_boolq_performance.md) | hard | | | 0 | | [ml_dev_bench_channel_vit_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_channel_vit_implementation.md) | hard | | | 0 | | [ml_dev_bench_channel_vit_implementation_easy](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_channel_vit_implementation_easy.md) | hard | | | 0 | | [ml_dev_bench_channel_vit_implementation_no_test](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_channel_vit_implementation_no_test.md) | hard | | | 0 | | [ml_dev_bench_cifar100_performance](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_cifar100_performance.md) | hard | | | 0 | | [ml_dev_bench_cifar_10_lt_performance](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_cifar_10_lt_performance.md) | hard | | | 0 | | [ml_dev_bench_dataset_not_available_download](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_dataset_not_available_download.md) | medium | | | 0 | | [ml_dev_bench_dataset_preprocess](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_dataset_preprocess.md) | hard | | | 0 | | [ml_dev_bench_full_train_workflow_performance_test](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_full_train_workflow_performance_test.md) | hard | | | 0 | | [ml_dev_bench_full_train_workflow_setup_test](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_full_train_workflow_setup_test.md) | hard | | | 0 | | [ml_dev_bench_improve_cifar10_baseline](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_improve_cifar10_baseline.md) | medium | | | 0 | | [ml_dev_bench_improve_segmentation_baseline](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_improve_segmentation_baseline.md) | hard | | | 0 | | [ml_dev_bench_lora_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_lora_implementation.md) | hard | | | 0 | | [ml_dev_bench_mcts_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_mcts_implementation.md) | hard | | | 0 | | [ml_dev_bench_mla_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_mla_implementation.md) | hard | | | 0 | | [ml_dev_bench_mla_implementation_hidden_tests](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_mla_implementation_hidden_tests.md) | hard | | | 0 | | [ml_dev_bench_nan_loss_debug](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_nan_loss_debug.md) | medium | | | 0 | | [ml_dev_bench_noisy_dataset_download](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_noisy_dataset_download.md) | hard | | | 0 | | [ml_dev_bench_noisy_label_annotation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_noisy_label_annotation.md) | hard | | | 0 | | [ml_dev_bench_normalization_bug](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_normalization_bug.md) | medium | | | 0 | | [ml_dev_bench_parse_logs](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_parse_logs.md) | medium | | | 0 | | [ml_dev_bench_ppo_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_ppo_implementation.md) | hard | | | 0 | | [ml_dev_bench_pretrained_bert_base_uncased_load](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_pretrained_bert_base_uncased_load.md) | medium | | | 0 | | [ml_dev_bench_pretrained_model_load_from_torchvision](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_pretrained_model_load_from_torchvision.md) | medium | | | 0 | | [ml_dev_bench_shape_mismatch_output](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_shape_mismatch_output.md) | medium | | | 0 | | [ml_dev_bench_shape_mismatch_train](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_shape_mismatch_train.md) | medium | | | 0 | | [ml_dev_bench_small_dataset_overfit](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_small_dataset_overfit.md) | hard | | | 0 | | [ml_dev_bench_training_files_debug](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_training_files_debug.md) | medium | | | 0 | | [ml_dev_bench_var_implementation](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_var_implementation.md) | hard | | | 0 | | [ml_dev_bench_vit_debugging](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_vit_debugging.md) | medium | | | 0 | | [ml_dev_bench_wandb_logging](https://harnessreport.com/tasks/ml_dev_bench/ml_dev_bench_wandb_logging.md) | medium | | | 0 | --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp