SpecEnvoy
EN·中文

‹ Blog

What's inside StarlingType

Published 7 June 2026

A plain-language look at how StarlingType turns your speech into text — in English and Chinese — without sending a single second of audio to the cloud.

You press nothing. You say your wake word, then talk — switching between English and Chinese mid-sentence if you like — and the words appear right where your cursor is. No menus, no "switch language" button, no copy-paste.

It feels a little like magic. It isn't. It's three pieces of genuinely excellent speech technology, all running on your own device. Here's what's actually happening, in human terms.

The one promise that shapes everything: your audio stays put

Most voice tools send your recordings to a company's servers to be transcribed. StarlingType doesn't. Every step happens locally, on your Mac or phone. Your voice is never uploaded, never stored on a server, never used to train anyone's AI. There's no "cloud mode" to accidentally leave on, because there's no cloud at all.

That's not a slogan bolted on at the end — it's the constraint we designed the whole product around. It's why we chose the specific technology we did.

So how does it work? Three steps.

1. It listens for your wake word — efficiently. A tiny, always-on "listener" waits for phrases like "Hey Type" or 嘿小听. It's deliberately small and lightweight so it can listen without draining your battery or paying attention to anything else — it's only watching for those few wake phrases. We actually run two side by side, one tuned for Chinese and one for English, so either language wakes it up.

2. It understands what you said. Once it hears the wake word, it captures your sentence and transcribes it. This is where the bilingual magic lives: it figures out which language you're speaking on its own, sentence by sentence — so "send him the 合同 by Friday" comes out right, mixed languages and all. No setting to flip.

3. It types for you. The finished text drops straight in at your cursor (or onto your clipboard, your choice). Start to finish, it's fast enough to feel instant.

Why "bilingual" is the hard part — and why ours is good at it

Lots of dictation tools speak many languages if you tell them which one first. Real bilingual people don't talk that way — we slide between English and Chinese in the same breath, and most tools choke on that.

StarlingType handles it because the transcription engine was built from the ground up as a multilingual model — it's especially strong at Chinese and Cantonese — and it detects the language as it goes instead of making you pick. That's the whole reason the app exists.

Built on the best open speech tech in the world

We didn't invent speech recognition from scratch — almost nobody should. Instead, StarlingType stands on the shoulders of two of the most respected open-source speech projects anywhere:

Both are free and openly licensed (Apache-2.0), which is exactly why we can run them entirely on your device and give the app away for free.

Honest about what it is — and isn't

StarlingType is a dictation tool: it shines at speaking a sentence or a paragraph and having it typed for you, in a reasonably quiet spot. It's not a meeting transcriber, it won't label who-said-what across a room, and like every voice tool it can stumble on heavy background noise or unusual names. We'd rather it do one thing really well than ten things poorly.

Free, private, and yours

No account. No subscription. No audio leaving your device. Just your voice, in two languages, turned into text wherever you're typing.

Dictate in English and Chinese — without switching languages.

Try StarlingType →
Curious about the actual models — exact names, who trained them, and how they stack up against 2026's state-of-the-art? Read the under-the-hood technical write-up (in English).