Voice input for ChatGPT on Mac.
The prompts that get good answers are three or four sentences of context you already have in your head. Typing is the only reason they end up as six words.
Short prompts are a typing problem
Everyone has been told that better context produces better answers, and everyone still sends six words. Not because the advice is wrong, but because the context is the boring part to type: what you already tried, the constraint that rules out the obvious approach, the audience, the format you want back.
You know all of it before you start. Writing it down is transcription, and transcription is where speech is about three times faster. Say the whole thing and you get the prompt you were told to write rather than the one your hands were willing to produce.
Typed, not spoken to the app
Hold it to talk, tap it to toggle, Escape cancels. Pick the key in Settings.
There is a real difference between dictation and talking to an assistant. Dictation puts text in the box and stops. You can read it back, cut the sentence that came out wrong, add the thing you forgot, and only then send it. That editing pass is where prompts actually get good, and it disappears the moment the audio goes straight through.
- Put the cursor in the prompt box, in the app or in a browser tab.
- Hold your shortcut, which you chose in Settings during setup.
- Talk through the whole thing. There is no session limit, so context is one press.
- Let go. The text lands in about a fifth of a second, punctuated.
- Read it, fix it, send it.
The same shortcut everywhere else
Prompting is not a separate workflow with its own tooling. The message you send after reading the answer, the commit that follows, the note explaining what you learned, the email to a colleague: those are all in different applications, and one shortcut covers them because Zumbo types where the cursor is rather than integrating with anything.
Getting the specifics right
Prompts are dense with proper nouns: libraries, services, file names, internal systems, the name of the thing you are building. Those are exactly what a general speech model spells as ordinary English, and a prompt with three wrong names will get you a confident answer about the wrong thing. Turn on the developer tools vocabulary pack, then correct anything else once with Teach, and it is heard your way from then on.
What Zumbo does with the audio
Nothing leaves your Mac. There is no account, no API key and no network request needed to transcribe, because the speech engine is inside the app. The transcription and the sending are two separate decisions, and only the second one involves anybody else.
Related: dictating into Claude Code, dictation for developers, and how Zumbo keeps your voice on your Mac.
Does ChatGPT have voice input already?
It has its own voice features and they change often enough that this page will not describe them. What system wide dictation adds is one shortcut that behaves identically in the ChatGPT window, the browser tab, your editor and your email, rather than a different gesture per app.
Why does the prompt need to be long?
Because the answer is only as good as the context. The difference between a useful reply and a generic one is usually three sentences you already know and did not type.
Does my prompt get transcribed in the cloud?
Not by Zumbo. The speech engine ships inside the app and composes on your Mac, then types the finished text into the box. What you choose to send afterwards is your decision, made with the words in front of you.
Will it spell the technical parts right?
Turn on the developer tools vocabulary pack for the common ones, and correct anything else once with Teach. The corrections stay on your Mac.
Can I edit before sending?
That is the point of typing rather than talking to the app directly. The text lands in the box and stops. You read it, change it and send it when you are ready.
Free for 3 days, everything unlocked. Then $15 once. macOS 14 or later, Apple silicon.