Overmind: Turn Your Agent's Production Traces Into Models You Own
The reflexive move in agent-building is to reach for the biggest model. Overmind is a bet on the opposite instinct: the best model for an agent is often the small one trained on exactly the task it does all day — labeling emails, calling a specific API, routing a ticket. Overmind is the open-source stack that turns that idea into a workflow.
What it does
Point Overmind at your agent’s codebase and it turns production traces (or any dataset) into a fine-tuned model, benchmarked against the eval metrics you define, and served through one unified API. No ML infrastructure to stand up. Crucially: the weights are yours — download, retrain, or roll back at will.
It’s less a single tool than a closed loop with six stages:
- Agent & Capabilities — a graph of your agent (prompts, tools, tasks) scanned straight from the repo
- Observability — OpenTelemetry traces, scored as they arrive and matched to the capability that produced them
- Datasets — production traces or uploaded files become versioned eval/training datasets
- Eval — a per-capability definition of “good,” measured on live traces and in batch
- Optimisers — prompt, tool, and control-flow experiments in your repo; the winner lands as a git diff
- Models — fine-tunes you own, benchmarked against production, served alongside frontier models on one OpenAI-compatible API
The workflow is agent-native
What makes it feel modern is that you drive it from inside your coding agent. Sign up, pick Cursor / Claude Code / OpenCode / Codex, and paste the onboarding prompt with your agent’s repo open. It installs Overmind, runs init and sync, and builds the context graph. After that it’s all slash commands in the same chat:
| Command | What it does |
|---|---|
/overmind ensure-tracing | Inspect traces and instrument the agent |
/overmind dataset | Build, clean, upload or export a dataset |
/overmind finetune | Fine-tune, deploy and smoke-test a model |
/overmind optimise | Run prompt and code optimisation |
/overmind backtest | Compare models against the agent’s own traces |
That backtest step is the quiet killer feature — you’re not comparing models on generic benchmarks, you’re comparing them on your agent’s own historical traces. “Good” means good at your job, not someone’s leaderboard.
How the training actually works
The marketing glosses over the mechanics, so here’s what’s under the hood. There’s no magic trainer and it isn’t Unsloth — it’s plain supervised fine-tuning (SFT) on GPUs, orchestrated for you:
- Compute runs on Modal. You supply
MODAL_TOKEN_ID/MODAL_TOKEN_SECRETand Overmind submits jobs to a Modal SFT worker (modal_sft_worker.py) for training and a vLLM worker (modal_vllm_worker.py) for serving. Checkpoints land in S3; gated base weights come from Hugging Face. Unsloth appears only as one base model in the catalogue (unsloth/Muse-Glimmer-30B), not as the training engine. - LoRA or full fine-tune — and often the model decides for you. Small dense models (Qwen 3/3.5, Llama 3.x, Gemma, Liquid LFM2.5, Cisco Antares) can do full fine-tunes. Larger dense models (70B) and every mixture-of-experts model (GPT-OSS 20B, Qwen Coder MoE, Nemotron Lightning) are LoRA-adapter only — the picker just tells you, no choice offered. Over the MCP API you force it with
"training_type": {"type": "Lora"}or{"type": "Full"}. - Model-specific preprocessing, not label repair. After you launch, CPU preprocessing runs on Modal using the matching tokenizer, chat template, and loss mask for the chosen base. Rows that overflow the training-method’s context limit, have invalid tool formatting, or lack supervised next-token targets are blocked, never silently truncated. The GPU then consumes the validated token artifact.
- Hyperparameters are per-experiment. One run can launch several experiments side by side (two learning rates, two base models) each with its own Epochs / Learning rate / Batch, plus Rank / Alpha / Dropout for LoRA runs. You watch per-step train+validation loss, LR schedule, token accuracy, and gradient norm live, with an ETA and credits-used counter.
- Benchmarked before it ships. The same eval set and frozen judges score baseline-vs-trained (and optionally your codebase incumbent) on held-out data. For label/classification tasks you also get per-class precision/recall/F1 and a confusion matrix — not just “it improved.”
The honest summary: Overmind is the control plane (dataset contracts, tokenization, LoRA/full selection, eval harness, serving), and Modal is the GPU plane. You own the resulting weights — downloadable from the model page or via overmind model download-checkpoint.
How to use it (sign-up to served model)
Option A — Hosted (fastest):
- Sign up at console.overmindlab.ai.
- On Get started, pick your coding agent — Cursor, Claude Code, OpenCode, or Codex — and paste the onboarding prompt into it with your agent’s repo open. That installs the
overmindCLI, runsovermind init+overmind sync, and builds the capability graph. - Instrument tracing:
pip install "overmind[tracing]", setOVERMIND_API_KEY(project key from Console → Settings), and run/overmind ensure-tracing.providers="auto"wraps the LLM SDKs you already use over OpenTelemetry — no key means tracing is simply off and nothing breaks. (Already using Langfuse / LangSmith / Braintrust / Galileo? Sync them through a connector instead.) - Let real traffic flow, then
/overmind datasetto turn traces into a versioned train/eval dataset. /overmind finetune— it checks readiness, estimates cost, launches the run, and follows it to completion.- Make live to switch your capability alias
overmind/<capability-uuid>to the new model, or Copy prompt to have your coding agent wire your LLM client to it via the OpenAI-compatible endpoint.
Option B — Self-host (data stays in your network):
git clone https://github.com/overmind-core/overmind.git && cd overmind
cp .env.example .env # set required keys (below)
docker compose up -d # Postgres, Redis, API :8000, Console :5173, Celery workers, Grafana :3001
First boot runs migrations, seeds the built-in evaluators, and exposes Swagger at /api/docs/. Sign in at http://localhost:5173 with any email/password; docker compose exec api python manage.py seed_demo --owner <you> loads a full demo workspace. The API refuses to start until every required key is set and names the missing one:
| Key | Used for |
|---|---|
OPENROUTER_API_KEY | Judges, evals, trace scoring, optimiser scoring, frontier inference |
MODAL_TOKEN_ID / MODAL_TOKEN_SECRET | Modal training + serving workers |
INFERENCE_API_URL / INFERENCE_API_KEY | Serving trained models (endpoint from modal deploy + shared secret) |
AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_BUCKET_NAME | Fine-tuning checkpoint archive (S3) |
HF_TOKEN | Gated Hugging Face base models (also set in the Modal secret) |
Fine-tuning and serving also need the Modal workers deployed (modal deploy overbae/modal/modal_vllm_worker.py, register_model.py, modal_sft_worker.py) and a Modal secret named overmind-inference holding the AWS keys, INFERENCE_API_KEY, and HF_TOKEN. The platform (overbae/, frontend/) is AGPL-3.0; the SDK/CLI (overmind/) is MIT.
Why it matters
The industry spent two years making agents by wrapping ever-larger general models. Overmind represents the maturing counter-current: specialization as infrastructure. For developers who want tight, repeatable control over narrow agent workflows — the email-labeler, the invoice-classifier, the tool-router — this is a clean path from “we have traces” to “we have a small model we own that’s measurably better at this one thing.” And it closes the loop: deploy, trace, evaluate, retrain, repeat.
Repo: github.com/overmind-core/overmind · Site: overmindlab.ai