A 60M model formats your dictation in under 100ms. A 14M model understands “next Tuesday at 6.” We tune them, then build the software around them.
One purchase, lifetime updates. No subscriptions, no accounts.
If a small model on your Mac beats the cloud, that's what we ship.
The places agents work well — and the layer that makes them faster.
Windows, iOS and Android are on the way.
Why on your machine
For years, AI has meant renting a giant model behind an API. We think most everyday intelligence belongs on the machine in front of you — faster, and private. And where a big cloud model is the right tool — your agent — our apps work with the one you already have.
dictation formatted by a 60M model, on-device
parameters to turn plain requests into structured tool calls
parameters to resolve “next Tuesday at 6pm” deterministically
Measured, not promised — published at lab.haku.sh
From model to app
We start from the best small local models, train them further for one job each, and quantize them to run fast on Apple silicon.
Instantly, offline, privately.
Apps you use daily and own outright.
The apps
Two shipping now, two on the way. macOS 14+ · no subscriptions · every app a one-time purchase.
dictation at the speed of thought.
The local intelligenceTranscription plus a 60M formatting model, fully on-device. Your voice never uploads.
where AI agents work as a team.
The local intelligenceSemantic history and code search on local embeddings. You stop being the messenger.
the web, finally usable by your agent.
The local intelligenceAn on-device ranker hands your agent the right controls — search, fill, book, buy.
your code, searched by meaning.
The local intelligenceOn-device code embeddings, indexed locally. Your code stays on your disk.
The research
We publish the results either way — what works graduates into the apps.