{"harness": {"recipes": 3, "summary": "Codex CLI (openai/codex) is a Rust coding agent. Its CLI is the codex-rs workspace crate codex-cli (binary `codex`); the npm package in codex-cli/ only wraps prebuilt binaries. The non-interactive entrypoint is `codex exec [OPTIONS] -- \"<prompt>\"`, which runs one headless session. `--json` streams JSONL events, `-o FILE` writes the final message, `--dangerously-bypass-approvals-and-sandbox` turns off approvals and the sandbox, `--skip-git-repo-check` allows non-git dirs, and CODEX_HOME holds the config and the rollout trajectories. Compiling the workspace needs more than 8 GB of RAM, and the previous build was OOM-killed. The overlay therefore installs the official statically linked musl `codex` binary that upstream CI builds from this workspace (GitHub release rust-v0.156.1, x86_64 or aarch64 chosen by `uname -m`). The repo source is still COPYed to /opt/harness/src, and `--build-arg CODEX_INSTALL=source` compiles from source on a large builder instead. Codex calls models only through the OpenAI Responses API (`POST <base_url>/responses`, SSE); `wire_api=\"chat\"` is rejected. So the run command starts a stdlib-Python shim on 127.0.0.1:8790 using a uv-managed Python under /opt/harness. The shim translates Responses requests into streaming chat-completions calls to $PROXY_URL/v1/chat/completions and converts the results back. Function tools map 1:1, and freeform custom tools such as apply_patch become functions with a single `input` string. Codex reaches the shim through a custom provider in CODEX_HOME/config.toml, with model gpt-5.5.", "recipe": "openai-codex@55543d87724bb66bdd51bde65254feb9b4c9ed10", "tasks_passed": 4, "harness": "openai-codex", "first_run": "2026-09-23T23:05:28", "domains": ["swe"], "runs": 7, "base_image": "python:3.12-slim-bookworm", "finished": 7, "last_run": "2026-09-25T16:10:25", "repo": "https://github.com/openai/codex", "passes": 4, "last_run_id": "20260925T090258-openai-codex-polyglot_go_palindrome-products", "use_case": "Helps developers write, debug, test, and refactor code in any programming language from the terminal", "tasksets": ["aider_polyglot", "bird-bench", "humanevalfix", "swebench-verified"], "results": {"bird-bench/card_games__372": {"passes": 1, "last": "2026-09-24T22:56:01", "last_run": "20260924T225535-sw-codex-card_games__372", "last_tests": null, "last_outcome": "pass", "last_reward": 1, "runs": 1}, "aider_polyglot/polyglot_python_bowling": {"passes": 1, "last": "2026-09-23T23:05:28", "last_run": "20260923T222433-openai-codex-polyglot_python_bowling", "last_tests": {"summary": "31 passed", "total": 31, "passed": 31, "failed": 0, "agent_written": 0, "failed_names": []}, "last_outcome": "pass", "last_reward": 1, "runs": 1}, "aider_polyglot/polyglot_go_palindrome-products": {"passes": 0, "last": "2026-09-25T16:10:25", "last_run": "20260925T090258-openai-codex-polyglot_go_palindrome-products", "last_tests": null, "last_outcome": "fail", "last_reward": 0, "runs": 2}, "swebench-verified/psf__requests-5414": {"passes": 1, "last": "2026-09-24T23:04:32", "last_run": "20260924T230425-sw-codex-psf__requests-5414", "last_tests": {"summary": "131 passed, 1 xfailed, 158 errors", "total": 0, "passed": 0, "failed": 0, "agent_written": 0, "failed_names": []}, "last_outcome": "pass", "last_reward": 1, "runs": 2}, "humanevalfix/python-12": {"passes": 1, "last": "2026-09-24T22:55:05", "last_run": "20260924T225434-sw-codex-python-12", "last_tests": {"summary": "1 passed", "total": 1, "passed": 1, "failed": 0, "agent_written": 0, "failed_names": []}, "last_outcome": "pass", "last_reward": 1, "runs": 1}}, "models": ["bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0"], "tasks_tried": 5, "api_style": "openai", "commit": "55543d87724bb66bdd51bde65254feb9b4c9ed10"}, "profile": {"evidence": "Based on README.md:1 (\"coding agent from OpenAI that runs locally\"), codex-rs/protocol/src/prompts/base_instructions/default.md:1-7 (system prompt describes capabilities: shell commands, apply patches, planning), codex-rs/core/src/tools/handlers/mod.rs:1-36 (tool handlers: apply_patch, unified_exec/shell, view_image, plan, multi_agents, web_search), codex-rs/tools/src/lib.rs (tool definitions for file editing and execution), and test suites in codex-rs/core/tests/suite/ showing Git operations, file editing, shell execution, and multi-agent coordination.", "harness": "openai-codex", "domains": ["swe"], "source": "claude -p", "cost_usd": 0.47110425000000006, "capabilities": ["edits-files", "runs-shell", "runs-tests", "uses-git", "multi-agent", "long-horizon", "reads-docs", "browses-web"], "seconds": 97, "languages": [], "at": "2026-09-25T16:13:28", "use_case": "Helps developers write, debug, test, and refactor code in any programming language from the terminal", "not_for": ["Native GUI automation (CLI/TUI only)", "Real-time voice interaction (text-based)", "Database querying without shell access"], "commit": "55543d87724bb66bdd51bde65254feb9b4c9ed10"}, "recommendations": {"model": "haiku", "at": "2026-09-25T16:14:10", "profile_source": "claude -p", "harness": "openai-codex", "recs": [{"score": null, "taskset": "aider_polyglot", "language": "cpp", "why": "Test whether the harness can handle compiled languages and their build/test toolchains, covering the 'any programming language' claim.", "task": "polyglot_cpp_allergies", "domain": "swe"}, {"score": null, "why": "Validate Rust support and assess if it handles the more complex state-tracking logic at medium difficulty (same task type it passed in Python).", "domain": "swe", "language": "rust", "taskset": "aider_polyglot", "task": "polyglot_rust_bowling"}, {"domain": "swe", "language": "javascript", "why": "Test an interpreted language and its test ecosystem, rounding out language coverage across compiled/interpreted paradigms.", "task": "polyglot_javascript_triangle", "taskset": "aider_polyglot", "score": null}, {"taskset": "swebench-verified", "why": "Validate real-world debugging on a different codebase (documentation tool vs. HTTP library) to confirm the harness generalizes beyond its first passing real-world fix.", "score": null, "language": null, "domain": "swe", "task": "sphinx-doc__sphinx-8595"}, {"taskset": "usaco", "score": null, "domain": "swe", "language": "python", "task": "1090", "why": "Step up to algorithmic problem-solving where code must be written from scratch without a template, testing the 'write' part of the use case at harder difficulty."}], "source": "llm", "based_on_run": "20260925T090258-openai-codex-polyglot_go_palindrome-products", "cost_usd": 0.029698}, "runs": [{"run": "20260925T090258-openai-codex-polyglot_go_palindrome-products", "started": "2026-09-25T16:10:25", "finished": "2026-09-25T16:11:20", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "aider_polyglot", "name": "polyglot_go_palindrome-products"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 0, "verifier_rc": 0, "tests": null, "calls": 14, "seconds": 47, "input_tokens": 165125, "output_tokens": 3405, "errors": 0, "last_action": "exec_command: cd /app && rm palindrome_test.go", "outcome": "scored", "verifier_says": "reward 0"}, {"run": "20260924T230425-sw-codex-psf__requests-5414", "started": "2026-09-24T23:04:32", "finished": "2026-09-24T23:08:45", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "swebench-verified", "name": "psf__requests-5414"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 1, "verifier_rc": 0, "tests": {"summary": "131 passed, 1 xfailed, 158 errors", "total": 0, "passed": 0, "failed": 0, "agent_written": 0, "failed_names": []}, "calls": 92, "seconds": 240, "input_tokens": 2667787, "output_tokens": 17848, "errors": 0, "last_action": "exec_command: cd /testbed && python3 -m pytest tests/test_requests.py::TestPreparingURLs::test_preparing_bad_url -v | grep \"example.co\u2026", "outcome": "scored", "verifier_says": "reward 1 \u00b7 131 passed, 1 xfailed, 158 errors"}, {"run": "20260924T225535-sw-codex-card_games__372", "started": "2026-09-24T22:56:01", "finished": "2026-09-24T22:56:16", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "bird-bench", "name": "card_games__372"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 1, "verifier_rc": 0, "tests": null, "calls": 5, "seconds": 10, "input_tokens": 53112, "output_tokens": 419, "errors": 0, "last_action": "exec_command: cat > /app/answer.sql << 'EOF' SELECT COUNT(*) FROM cards WHERE toughness = '99'; EOF", "outcome": "scored", "verifier_says": "reward 1"}, {"run": "20260924T225434-sw-codex-python-12", "started": "2026-09-24T22:55:05", "finished": "2026-09-24T22:55:30", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "humanevalfix", "name": "python-12"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 1, "verifier_rc": 0, "tests": {"summary": "1 passed", "total": 1, "passed": 1, "failed": 0, "agent_written": 0, "failed_names": []}, "calls": 10, "seconds": 21, "input_tokens": 103230, "output_tokens": 1259, "errors": 0, "last_action": "exec_command: cd /workspace && python Python__12.py", "outcome": "scored", "verifier_says": "reward 1 \u00b7 1 passed"}, {"run": "20260924T224755-sw-codex-psf__requests-5414", "started": "2026-09-24T22:48:31", "finished": "2026-09-24T22:54:27", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "swebench-verified", "name": "psf__requests-5414"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 0, "verifier_rc": 1, "tests": {"summary": "15 failed, 116 passed, 1 xfailed, 158 errors", "total": 0, "passed": 0, "failed": 0, "agent_written": 0, "failed_names": []}, "calls": 69, "seconds": 181, "input_tokens": 1522341, "output_tokens": 15812, "errors": 0, "last_action": "exec_command: cd /testbed && python3 << 'EOF' \"\"\" Final comprehensive test of the fix. \"\"\" import sys import requests from requests.ex\u2026", "outcome": "scored", "verifier_says": "reward 0 \u00b7 15 failed, 116 passed, 1 xfailed, 158 errors \u00b7 verifier exited 1"}, {"run": "20260924T224300-sw-codex-polyglot_go_palindrome-products", "started": "2026-09-24T22:45:48", "finished": "2026-09-24T22:47:45", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "aider_polyglot", "name": "polyglot_go_palindrome-products"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 0, "verifier_rc": 0, "tests": null, "calls": 23, "seconds": 107, "input_tokens": 310150, "output_tokens": 5729, "errors": 0, "last_action": "exec_command: cat /app/palindrome_products.go", "outcome": "scored", "verifier_says": "reward 0"}, {"run": "20260923T222433-openai-codex-polyglot_python_bowling", "started": "2026-09-23T23:05:28", "finished": "2026-09-23T23:08:16", "status": "done", "kind": "harbor", "harness": "openai-codex", "task": {"taskset": "aider_polyglot", "name": "polyglot_python_bowling"}, "model": "bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0", "reward": 1, "verifier_rc": 0, "tests": {"summary": "31 passed", "total": 31, "passed": 31, "failed": 0, "agent_written": 0, "failed_names": []}, "calls": 37, "seconds": 164, "input_tokens": 861810, "output_tokens": 21322, "errors": 0, "last_action": "exec_command: python3 << 'EOF' # 10 frames of (5, 4) pairs # Each frame scores 9 (just open frames, no bonus) # Total = 10 * 9 = 90 pr\u2026", "outcome": "scored", "verifier_says": "reward 1 \u00b7 31 passed"}]}