Magnitude: Run Your Agent on Local Models, Tuned to Your Exact Hardware

By Prahlad Menon 3 min read

The pitch is simple, and it’s one more teams are landing on: you already own the computer — now it can do more of the work. Magnitude is an open-source inference engine that runs the best local models for your hardware and plugs straight into the agent you already use.

Why that’s useful

  • Run models on your own computer. No cloud round-trip for inference.
  • Avoid per-token API charges for anything you run locally.
  • Work offline once everything’s downloaded — prompts, files, and models never leave your machine.
  • Switch models without changing your agent. The agent points at one endpoint; you swap what’s behind it.

Magnitude also unloads models when they’re idle or your memory gets tight, and shares prefix caches across concurrent sessions so parallel agents don’t grind each other to a halt. Performance depends on your hardware and the model you choose — but the whole point is that it’s your hardware.

The part that’s actually novel

Most local engines ship kernels precompiled for broad classes of hardware and hope they fit. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip. The claimed result against llama.cpp:

  • Up to 2x faster overall
  • 92% faster decode on Metal (Apple Silicon)
  • 19% faster on CUDA (NVIDIA)
  • ~27% less memory per agent, reclaimed when agents stop

They get there by hand-writing optimized kernels for the most popular open-weight model families (Qwen, DeepSeek, Gemma, and friends) rather than staying fully generalist — which is how a specialized engine beats a general one.

How you use it

  1. Download Magnitude for macOS, Windows, or Linux, install it, and open the app. The desktop app bundles the magnitude CLI — no separate install.
  2. Pick a recommended model in Discover and download it. Magnitude profiles your machine and only suggests models that actually fit.
  3. Connect your agent in Connections and start working. One click wires up Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, or Cline. Anything else talks to it over the OpenAI-compatible API.

Why it matters

Local inference has always carried a tax: either you fight with a generalist runtime that leaves performance on the table, or you hand-tune kernels you don’t have time to maintain. Magnitude’s bet is that the tuning should happen automatically, on your box, for your silicon — and that the agent layer shouldn’t have to care which model is behind the endpoint. Free, private, Apache-2.0, and it works with the agent you’re already running. Worth a slot on any local-first setup.

Repo: github.com/magnitudedev/magnitude · Site: magnitude.dev