HarnessRouter: One API to Run Every Agent Harness (Codex, Claude Code, and More)
I write a lot about agents here, and there’s a recurring pain that never shows up in the demos: every agent harness has its own API, its own session model, its own streaming format, its own way of handling files and failures. Want to swap Claude Code for Codex? Rewrite your integration. Add DeepSeek Harness? Rewrite it again. HarnessRouter exists to kill that repetitive plumbing once and for all.
The idea in one line
HarnessRouter is the infrastructure layer that turns existing harnesses into plug-and-play backends. You build your agent product against one API; HarnessRouter runs tasks, retrieves results, and switches harnesses underneath — Codex, Claude Code, Hermes, PI, DSH, and more — without you touching the integration code each time.
It’s self-hosted, Apache-2.0, your keys, your infrastructure.
The Unified Harness Protocol (UHP)
The interesting bit isn’t the router — it’s the standard it implements. UHP is an open protocol that defines the common surface every agent harness actually needs:
- persistent sessions — stateful conversations/tasks that survive across calls
- streaming progress — live token/step updates in a common format
- files and artifacts — standardized handling of inputs and generated outputs
- cancellation — stop a running task cleanly
- structured failures — errors you can actually branch on, not stack-trace soup
And it exposes all of that through an OpenAI Responses-compatible API — so if your code already speaks that, adding agent harnesses is nearly free. There’s even a protocol conformance suite in the repo, which tells you this is meant as a genuine open standard other projects can implement, not a single-vendor lock-in.
Why this matters: harnesses as interchangeable parts
This is the same modularity story that made LLM gateways (LiteLLM, Portkey) valuable — but one level up the stack. A gateway lets you swap models behind one API. HarnessRouter lets you swap whole agent harnesses — the loop, the tool-calling, the session logic — behind one API.
Concretely, that unlocks:
- Hot-swapping frameworks without breaking production pipelines
- Head-to-head comparison — the repo benchmarks eight harness × model configurations on the same task, and notes the cheapest and fastest configs vary by task. Being able to measure that is the point.
- No rewrite tax — session tracking, streaming, files, and error recovery are solved once, in the layer, not re-implemented per tool.
Genuinely self-hostable, today
The Community Edition is one Docker command:
docker run -d --name harnessrouter \
-p 127.0.0.1:3000:3000 \
-v harnessrouter:/data \
harnessrouter/harnessrouter
You need Docker, ~4GB disk, and a provider API key — no account, no bundled model, no trial key. First launch installs the enabled harness CLIs; when the logs say ready on :3000, you open localhost:3000 and connect a provider. The named volume keeps your database, files, installed harness CLIs, and workspaces across restarts. (Default creds are harnessrouter/harnessrouter — change them immediately and keep it loopback-bound until you do.)
Where it fits in the bigger picture
This is the plumbing counterpart to the Paperclip story I also wrote about today. Paperclip orchestrates a team of agents like a company; HarnessRouter makes each agent’s runtime interchangeable underneath. Together they sketch the same future the whole ecosystem is converging on: agents as interchangeable, unified components — standard protocols, hot-swappable backends, your infrastructure, your keys — instead of a pile of bespoke, incompatible integrations.
For anyone building agent products, this is weeks of repetitive infrastructure work you get to just… not do.