Edge0: Running 35B Mixture-of-Experts Models in 3 GB on a Mac mini

By Prahlad Menon 4 min read

Sparse Mixture-of-Experts models made a promise that edge hardware never fully cashed: activate only a slice of the parameters per step, and you should be able to run a very large model cheaply. The catch is that the experts you aren’t using still have to live somewhere. If they all sit in RAM, sparse compute never turns into a smaller memory footprint — and on a laptop or Mac mini, memory is exactly what you don’t have.

Edge0 is an open-source streaming MoE inference framework built specifically for that problem. It fuses two ideas — SSD expert offloading and Prerouter routing prediction — into a single inference path, and runs today through MLX on Apple Silicon.

The core trick: the SSD is a memory tier

Edge0 doesn’t keep every expert weight resident. The SSD stores the complete model; memory only ever carries the immediate working set. For each step the router picks a handful of experts, Edge0 reads just those weights off disk, and keeps recently-used ones cached. The result is striking:

TierBase architectureWeight footprint (4-bit)Peak active memory
Edge0-35B-A3BQwen3.5-MoE 35B-A3B (40 layers / 256 experts)~23 GB~2.9 GiB
Edge0-8B-A1BLing 3.0 (24 layers / 128 experts)~4.2 GB~1.0 GiB

A 35-billion-parameter model with roughly 2.9 GiB of peak active memory. The total parameter count no longer equals the amount of memory that must stay resident at every step.

Why naive SSD offloading isn’t enough

If you just page experts from disk on demand, you put storage latency directly on the critical path for every generated token: compute, route, wait for the read, compute again — at every layer. That stall is what kills throughput.

Edge0 splits the path into two cooperating parts. The SSD holds the weights; a streaming expert pool manages caching and movement. The missing piece is knowing what to load before you need it — which is where the Prerouter comes in.

Prerouter: move the wait earlier

The Prerouter is a lightweight trained prediction head that uses hidden states and routing features to guess which experts are coming next. Its clever bit is a dual shift across layers and tokens: layer N emits a prediction the next layer uses for a later token. Predictions commit at generation-step boundaries, so expert loading overlaps with the forward pass instead of blocking it.

Swapping Prerouter in for native routing under the same model and workload yielded a maximum 59% decoding speedup. Gains vary with cache hits, routing width, and SSD speed — but because the predicted route is part of real execution, the quality numbers reflect the complete quantized-and-adapted pipeline, not an idealized one.

Edge0 also separates Prefill (prompt processing, full-layer batched compute) from Decode (token-by-token, cache- and prefetch-bound). Both tiers share the same streaming path, with routing width matched per model.

Measured on a Mac mini M4 Pro

Runtime tests used a 24 GB Mac mini M4 Pro, comparing the int4 Edge0 pipeline against the fp16 base models with OpenCompass:

Metricedge0-35bedge0-8b
Decode speed14.9–17.7 tok/s23.9–25.3 tok/s
Prefill (cold / warm)113 / 140 tok/s500 / 1428 tok/s
Peak active memory (short ctx)2.9 GiB1.0 GiB

Those are usable interactive speeds for a 35B model on a sub-$2k desktop — with the model itself streaming off the SSD.

Why it matters

Edge0 reframes the edge-LLM constraint. The question stops being “how much RAM does this model need?” and becomes “how fast is your SSD, and how good is your prefetch?” By making storage, data movement, and scheduling one system instead of three, it turns sparse activation into genuinely lower device requirements — the thing MoE always promised but rarely delivered on-device.

One recipe, every device: the project targets macOS, iOS, and Android, with Recover-LoRA adapters restoring quality lost to aggressive quantization. If you have Apple Silicon and an SSD, you can try it now.

Repo: github.com/Edge0-AI/edge0