What's inside StarlingType
Published 7 June 2026
A plain-language look at how StarlingType turns your speech into text — in English and Chinese — without sending a single second of audio to the cloud.
You press nothing. You say your wake word, then talk — switching between English and Chinese mid-sentence if you like — and the words appear right where your cursor is. No menus, no "switch language" button, no copy-paste.
It feels a little like magic. It isn't. It's three pieces of genuinely excellent speech technology, all running on your own device. Here's what's actually happening, in human terms.
The one promise that shapes everything: your audio stays put
Most voice tools send your recordings to a company's servers to be transcribed. StarlingType doesn't. Every step happens locally, on your Mac or phone. Your voice is never uploaded, never stored on a server, never used to train anyone's AI. There's no "cloud mode" to accidentally leave on, because there's no cloud at all.
That's not a slogan bolted on at the end — it's the constraint we designed the whole product around. It's why we chose the specific technology we did.
So how does it work? Three steps.
1. It listens for your wake word — efficiently. A tiny, always-on "listener" waits for phrases like "Hey Type" or 嘿小听. It's deliberately small and lightweight so it can listen without draining your battery or paying attention to anything else — it's only watching for those few wake phrases. We actually run two side by side, one tuned for Chinese and one for English, so either language wakes it up.
2. It understands what you said. Once it hears the wake word, it captures your sentence and transcribes it. This is where the bilingual magic lives: it figures out which language you're speaking on its own, sentence by sentence — so "send him the 合同 by Friday" comes out right, mixed languages and all. No setting to flip.
3. It types for you. The finished text drops straight in at your cursor (or onto your clipboard, your choice). Start to finish, it's fast enough to feel instant.
Why "bilingual" is the hard part — and why ours is good at it
Lots of dictation tools speak many languages if you tell them which one first. Real bilingual people don't talk that way — we slide between English and Chinese in the same breath, and most tools choke on that.
StarlingType handles it because the transcription engine was built from the ground up as a multilingual model — it's especially strong at Chinese and Cantonese — and it detects the language as it goes instead of making you pick. That's the whole reason the app exists.
Built on the best open speech tech in the world
We didn't invent speech recognition from scratch — almost nobody should. Instead, StarlingType stands on the shoulders of two of the most respected open-source speech projects anywhere:
- The "listening" and wake-word technology comes from the open-source Next-gen Kaldi project — the modern continuation of the work behind the influential Kaldi toolkit that a whole generation of speech systems was built on.
- The transcription engine is SenseVoice, an open multilingual model from Alibaba's speech-AI team, trained on hundreds of thousands of hours of audio.
Both are free and openly licensed (Apache-2.0), which is exactly why we can run them entirely on your device and give the app away for free.
Honest about what it is — and isn't
Free, private, and yours
No account. No subscription. No audio leaving your device. Just your voice, in two languages, turned into text wherever you're typing.
Dictate in English and Chinese — without switching languages.