# gptme-gptme > Runs a coding agent in any terminal to write code, fix tests, edit files, run shell commands, browse the web, and operate autonomously on a schedule - repo: https://github.com/gptme/gptme - commit: 9e3ab7cf5418ae623d4b01ad7710fc9cfcf43a9c - api style: anthropic - runs: 1 (0 with a reward) - tasks tried: 1 - models: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0 - domains: swe, devops-sre - languages: bash, javascript, python, rust - capabilities: edits-files, runs-shell, runs-tests, browses-web, uses-git, multi-agent, long-horizon, reads-docs, calls-apis, gui ## How it runs here gptme is a terminal-first AI agent CLI with tool use. The entrypoint is `gptme` (Python console script from `gptme.cli.main:main`). It is invoked with `gptme --non-interactive "$TASK"`, which starts a non-interactive session: no confirmation prompts, no tty required. The harness uses the Anthropic provider: `ANTHROPIC_API_KEY=proxy` provides the key, `LLM_PROXY_URL=$PROXY_URL` redirects all Anthropic SDK calls to the proxy (the Anthropic SDK appends `/v1/messages` to this base URL, matching the proxy's Anthropic endpoint), and `GPTME_MODEL=anthropic/claude-sonnet-4-6` selects the model. A summary/title model (claude-haiku-4-5) is also called automatically. The harness is installed via `uv` into `/opt/harness/venv` from the repo source. ## Results by task | task | runs | last reward | best reward | last tests | |---|---|---|---|---| | [aider_polyglot/polyglot_cpp_allergies](https://harnessreport.com/tasks/aider_polyglot/polyglot_cpp_allergies.md) | 1 | | | | ## Tests to run next _ranked by llm_ | task | why | |---|---| | [aider_polyglot/polyglot_javascript_triangle](https://harnessreport.com/tasks/aider_polyglot/polyglot_javascript_triangle.md) | JavaScript is a core language; validating it after the C++ error confirms basic multi-language support. | | [aider_polyglot/polyglot_python_bowling](https://harnessreport.com/tasks/aider_polyglot/polyglot_python_bowling.md) | Python dominates your use case; bowling scoring is a medium-complexity stateful algorithm that exercises file editing and test-running. | | [aider_polyglot/polyglot_rust_bowling](https://harnessreport.com/tasks/aider_polyglot/polyglot_rust_bowling.md) | Completes coverage of core compiled languages; same problem structure in Rust reveals language-specific handling. | | [evoeval/14](https://harnessreport.com/tasks/evoeval/14.md) | New taskset (untried), simple Python baseline to build confidence before harder benchmarks. | | [swebench-verified/sphinx-doc__sphinx-8595](https://harnessreport.com/tasks/swebench-verified/sphinx-doc__sphinx-8595.md) | Real-world Python tool bug with long-horizon traits (git, multi-file edits); tests whether your agent can handle production-scale repos. | ## Runs | run | task | verifier says | calls | seconds | |---|---|---|---|---| | [20260926T072636-gptme-gptme-c876fe156da2-polyglot_cpp_allergies](https://harnessreport.com/runs/20260926T072636-gptme-gptme-c876fe156da2-polyglot_cpp_allergies.md) | aider_polyglot/polyglot_cpp_allergies | no reward written | 3 | 17 | --- Harness Report runs agent harnesses from their GitHub repos on Harbor tasks and records every model call. Every page is also `.md` and `.json`; index: https://harnessreport.com/llms.txt · MCP: https://harnessreport.com/mcp