Laya-CoreML: A Local AI That Decides in ~5 Milliseconds on the Neural Engine
This is now the fourth post in a thread I didn’t plan to write: Kev (an open decision model), Laya (fast, local, calibrated), Keel (a decision model doing a real job inside a coding agent), and now Laya-CoreML — the same Laya model ported to run on Apple’s Neural Engine, deciding in about 5 milliseconds, fully offline.
If Laya was “fast, local decisions,” this is “fast, local decisions on the chip in your Mac — and here’s the honest power bill.”
What it is
Laya-CoreML takes the open-weight Laya typed-decision model — the one that answers choice / ordinal score / yes-no (“noul”) questions with calibrated probabilities in a single forward pass, no generated tokens — and ships validated ports to Core ML targeting the Apple Neural Engine (ANE).
The pitch: run a real AI decision agent on your Mac that answers in ~5ms, uses a fraction of the energy of standard setups, and never touches the network.
The Snake demo is the hook
The repo’s centerpiece is delightful: a real Laya model plays terminal Snake, itself, live. The screen shows, in real time:
- the model’s move probabilities (up / down / left / right)
- score, length, latency
- a visible cycle-safety layer and safety interventions
Across three uncapped 600-step episodes it sustained ~49–50 decisions/sec with zero deaths (two safety interventions total). It’s the clearest visualization I’ve seen of a System 1 model doing what it’s built for: making a bounded, typed choice, many times a second, with its confidence exposed.
pip install 'laya-coreml[demo]'
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
laya-coreml-snake --model ./models/snake
Download once, then play offline — no PyTorch, Transformers, or MLX needed for inference.
The benchmarks — and why the honesty matters
Here’s what makes this repo trustworthy. It does not oversell. The measured M3 Max numbers, one 91-token question:
| Metric | Compiled MLX FP16 | Core ML ANE FP16 | Core ML ANE W8 |
|---|---|---|---|
| P50 / P95 latency | 6.94 / 7.39 ms | 4.98 / 5.31 ms | 4.88 / 5.23 ms |
| Mean system power | 61.39 W | 30.75 W | 27.39 W |
| Energy / decision | 0.4288 J | 0.1540 J | 0.1344 J |
| Speed gain | 1× | 1.39× | 1.42× |
| Energy gain | 1× | 2.78× | 3.19× |
And then, in plain text: “the requested 10× improvement was not achieved.”
That line is worth pausing on. A lot of “local AI destroys the cloud” posts quietly bury the caveats. This one leads with them. The speed win over an already-fast MLX baseline is a modest ~1.4×. The real story is energy: ~2.8–3.2× less joules per decision, from moving work onto the Neural Engine instead of GPU. For anything battery-powered or always-on — an agent making thousands of routing/guardrail decisions — that’s the number that matters.
The calibration bug they caught (this is the good stuff)
Buried in the README is a detail that tells you these are serious people. Laya’s decisions are only useful if the probabilities are honest — that’s the whole point of the proper-scoring-rule training I described for Laya. But calibration temperatures can go wrong:
The shipped
choice:11+bucket temperature is 0.1006, which would sharpen logits ~10× and report a coin flip as near-certainty.
So they clamp fitted temperatures to [0.5, 5.0] before use, keep the raw values available for inspection (agent.temperature_raw), and emit a RuntimeWarning naming every clamped bucket at load. That’s exactly the kind of guardrail that separates “cool demo” from “thing you’d put in production” — and it rhymes with the calibration-temperature-fitting step in Laya’s own fine-tuning loop.
The honest limits
Also refreshingly stated up front:
- ANE bundle: 96-token total limit (question + options + state). Longer inputs error out — use
aac6fef/laya-multilingual-coremlfor the 1024-token general model. - Requires Apple Silicon, macOS 15+, Python 3.11–3.13.
- Snake’s decision-rate numbers include rendering serialization; single-question latencies exclude load/warmup. They’re careful to say which is which.
Why this closes the loop
Four posts in, the arc is clear:
- Kev — you can build an open decision model.
- Laya — you can make it fast, local, and calibrated.
- Keel — you can put it to work as a permission-checked router.
- Laya-CoreML — you can run it on the Neural Engine at ~5ms and ~⅓ the energy, offline, with the calibration bugs actually caught.
This is the same lesson as our RF-SRC memory-optimization work: small, purpose-built models that emit calibrated probabilities and run on hardware you already own beat a giant generalist for a huge class of structured decisions. Laya-CoreML just proves it down to the joule — and, notably, proves it without inflating the claim.
No cloud, no lag, no hype. Just a tiny model making honest decisions on your chip.