Speakr: Self-Hosted AI Transcription That Turns Recordings Into a Searchable, Private Knowledge Base
Most transcription tools ask you to upload your most sensitive conversations — client calls, legal consultations, research interviews, family memories — to someone else’s cloud. Speakr takes the opposite stance: it runs entirely on your own infrastructure, turning audio into organized, searchable, intelligent notes without a single byte leaving your control (unless you explicitly choose a cloud ASR connector).
It’s AGPL-3.0 licensed, Docker-deployed, and — refreshingly — genuinely feature-complete rather than a weekend demo.
Bring your own engine
The heart of Speakr is a connector-based architecture that auto-detects your transcription provider. The recommended one is self-hosted WhisperX, and for good reason: it’s the only backend that unlocks the full speaker feature set —
- Speaker diarization — automatic who-said-what labeling.
- Voice profiles — recognize the same person across different recordings via voice embeddings. That’s a genuinely advanced feature most transcription apps don’t touch.
- Custom vocabulary and hotwords — bias the transcriber toward names, jargon, and acronyms it would otherwise mishear, configurable globally, per tag/folder, or right in the upload dialog.
If you’d rather not run a GPU container, Speakr also connects to OpenAI (gpt-4o-transcribe-diarize), Mistral/Voxtral, VibeVoice (self-hosted via vLLM), MOSI/Mossland, AssemblyAI, OpenASR, and Alibaba FunASR. The right connector is picked automatically from your config.
From recording to knowledge
The pipeline goes well beyond raw transcripts:
- Flexible capture — record from your mic, system/browser-tab audio, or both mixed; or drag-and-drop existing files. It even surfaces virtual devices like BlackHole, VB-Cable, Voicemeeter, and Stereo Mix.
- Long sessions stream to the server during capture, so recordings can run for hours and survive a page reload.
- Hands-off intake — a watched “black hole” folder auto-imports and processes anything dropped into it.
- Summaries, event extraction, and per-recording chat — auto-generated summaries with fully customizable prompts, action-item/calendar-event extraction, and a floating chat panel whose answers cite transcript timestamps as clickable chips that jump playback to that exact moment.
Smart tags: the killer idea
Speakr’s tags aren’t just labels — each carries its own AI prompt and ASR settings, and they stack. This turns one recording into whatever you need:
- A “Recipe” tag converts messy narrated cooking into a formatted recipe with ingredients and numbered steps.
- A “Study Notes” tag turns a lecture into an organized outline of concepts and definitions.
- An “Action Items” tag ignores the discussion and returns just decisions, tasks, and deadlines.
And they layer: Recipe + Gluten Free yields a formatted recipe with substitution suggestions; Client Meeting + Legal Review surfaces requirements and legal implications. Order matters — format tags first, focus tags second.
Inquire Mode: an AI researcher over your whole library
This is the standout. Inquire Mode is semantic search and natural-language chat across your entire library at once. The opt-in agentic beta (ENABLE_INQUIRE_AGENT=true) upgrades it to an AI researcher that iteratively searches, lists, and reads recordings until it can answer — showing every step live and citing sources with numbered links that open the recording at the cited segment.
Ask “what did we decide about the pricing change, and did anyone disagree?” and it reads across the relevant meetings and cites each position back to where it was said. Crucially, you decide what it may read: transcripts always, summaries and private notes only if you allow them.
Built for teams, not just tinkerers
Speakr scales past a single user with a serious set of collaboration and governance features: multi-user SSO against any OIDC provider (Keycloak, Azure AD, Google, Auth0, Pocket ID), groups with auto-sharing tags, granular internal sharing, admin-controlled public links, retention policies with auto-deletion (7-day standups, 7-year legal retention, permanent architecture-decision archives), per-user usage budgets, a REST API v1 with Swagger UI for n8n/Zapier/Make automation, and HMAC-signed webhooks. It’s also a mobile-first installable PWA with offline support, light/dark themes, an incognito mode, and a UI in seven languages.
Getting started
mkdir speakr && cd speakr
wget https://raw.githubusercontent.com/murtaza-nasir/speakr/master/config/docker-compose.example.yml -O docker-compose.yml
wget https://raw.githubusercontent.com/murtaza-nasir/speakr/master/config/env.transcription.example -O .env
nano .env # add your API keys / ASR_BASE_URL
docker compose up -d
# Access at http://localhost:8899
There’s a lightweight learnedmachine/speakr:lite image (~725MB vs ~4.4GB) that skips PyTorch — every feature still works, only Inquire Mode’s semantic search falls back to basic text search.
The takeaway
Speakr is one of the most complete self-hosted AI tools I’ve seen lately: it treats your recordings as a private, queryable knowledge base rather than a pile of transcripts, and it does the hard parts — cross-recording voice identity, agentic citation-backed search, tag-driven prompt pipelines — while keeping everything on hardware you control. For anyone who handles sensitive audio and refuses to ship it to a third party, it’s an easy recommendation.
Repo: github.com/murtaza-nasir/speakr · Docs: murtaza-nasir.github.io/speakr