Private dictation for macOS. Press fn, speak, and your words land wherever you type,
transcribed on your Mac.
Free and open source. No account, no subscription.
Live typing: words appear as you speak, greyed while provisional, then confirmed at each pause.
| Tap fn to start and stop, or hold it to talk. The text is pasted into Word, Teams, VS Code, Rider, Mail or wherever the cursor is. | Whisper, Parakeet and Apple Speech run locally on Apple Silicon. A cloud engine is there only if you choose it. | Named prompts clean up, format or translate a dictation. Command mode rewrites selected text with a spoken instruction. |
A small HUD appears next to the caret while you speak. Tap fn again and the transcription is pasted in place. Esc cancels at any moment.
The HUD shows the microphone level while recording, a rainbow edge while transcribing and while an AI prompt runs, and folds away when it is done. It floats in Liquid Glass tinted with one of nine colours.
Dictating over a Finder window or a web page with no text field? The text stays in the HUD, still editable, with Copy and Close.
![]() History (opt-in): click to copy, ▶ to paste again. |
![]() Settings: models, hotkeys, HUD, languages, vocabulary, voice macros and AI providers. |
| Model | Runs | Good for |
|---|---|---|
| Whisper Large v3 Turbo | On device · ~1.6 GB | Best quality, many languages |
| Whisper Large v3 | On device · 626 MB | Recommended by Argmax, smaller download |
| Whisper Small | On device · ~200 MB | Quick tests |
| Parakeet v3 | On device · ~0.6 GB | Very fast, 25 European languages |
| Parakeet Ultra | On device · ~0.6 GB | Parakeet v3 post-trained for accuracy, same speed |
| Apple Speech | Built into macOS | No download |
| Gemini API | Cloud · your API key | Optional; the recording is sent to Google |
A silence and hallucination filter (Silero VAD) keeps invented phrases like “Thank you.” out of your text.
- Download the latest
SimpleWhisper-x.y.z.zipfrom Releases, unzip it and move SimpleWhisper to Applications. - The app is not notarized yet, so clear the quarantine flag once:
xattr -dr com.apple.quarantine /Applications/SimpleWhisper.app
- Open it and allow Microphone, Accessibility and Input Monitoring when asked.
- Press fn and start talking. Turn on automatic updates in Settings › General › Updates.
The full reference, folded into sections. Click a section to open it.
Hotkeys
| Key | Action |
|---|---|
| fn (tap) | Start / stop dictation |
| fn (hold, > 0.4 s) | Push-to-talk: release to stop |
| Esc | Cancel recording or processing |
| Letter / Space | While recording: pick the prompt with that shortcut / plain text |
| Control | While recording: run the dictation as a command on the selection |
| Control + letter / Space | In Live typing: pick a prompt / plain text (letters type into the editor) |
| fn + any key | A keyboard shortcut: the recording is cancelled silently |
The dictation key can also be left/right Command or left/right Option (Settings › General › Hotkey). Optional double-press mode: two quick presses toggle, press-release-hold is push-to-talk, a single press does nothing. Pressing the dictation key together with another modifier (Control, Shift, Option, Command) never starts a recording.
Live typing
Settings › General › Output, or the menu bar. An editor opens under the HUD and fills with text as you speak: the utterance in progress is re-transcribed about once a second and shown greyed, and it is confirmed at the next pause (about 0.7 s) or after 25 s.
- Click anywhere to move the caret and keep dictating there, select text to replace it, or edit with the keyboard; the editor grows with the text.
- The clipboard button in the corner pastes the clipboard at the caret.
- Stop dictation (fn, or release fn in push-to-talk) to run the selected prompt and insert the result into the original app. Esc discards it (kept in History when enabled).
- While the editor is open, letters and space type into it; Control + a prompt's letter picks that prompt and Control + Space plain text.
- The Gemini API engine has no live preview (every call is billed); its text appears at pauses.
Prompts, command mode and AI providers
- Prompts: named AI prompts (Clean up, Formal email, Bullet points, Translate to English, your own) run through Claude Code CLI (
claude -p), Codex CLI, Gemini CLI, Agy, a custom shell command, the Claude API, the OpenAI API, the Google Gemini API (keys in the Keychain) or Apple Intelligence. Clipboard and macro content is protected by placeholders so the AI never rewrites it. - Prompt shortcuts: each prompt can have a letter; press it while recording to use that prompt, press Space for plain text. Clicking the HUD opens the prompt menu too.
- Command mode (Settings › General): select text in the editor, start dictation, say what to do with it (“convert to markdown”, “translate to English”, “make it shorter”) and click the round ▶ button in the HUD or press Control (while still holding fn, or after a short fn press). The selection is read (Accessibility, or ⌘C), sent to the AI command with your instruction, and the result replaces the selection. With nothing selected the dictation is a question to the assistant, and the answer opens as a Markdown document.
- Settings › AI has two parts, prompts and command mode/assistant, each with its own provider.
Languages, vocabulary and voice macros
- Language: auto-detect among a configurable set of languages (default English; any of Whisper's 99 languages can be added), a fixed language, or auto-detect anything. Mixed sentences are fine.
- Vocabulary: words the models tend to get wrong plus aliases that are always corrected in the final text.
- Voice macros: say “schowek” / “clipboard” to insert the clipboard content (captured when recording starts); “nowa linia” / “new line” for a line break; custom text macros.
Spoken punctuation
Macros of type Punctuation, Polish and English: “przecinek”/“comma”, “kropka”/“period”, “znak zapytania”, “wykrzyknik”, “dwukropek”, “średnik”, “myślnik”, “cudzysłów”, “nawias otwarty/zamknięty”, “nowy akapit”/“new paragraph”. Marks are placed with proper spacing, the next sentence is capitalised, and punctuation the model already inserted is merged rather than doubled. One switch in Settings › Macros turns it off when those words are meant literally.
HUD, sounds, history and the result card
- Menu bar only (no Dock icon). The HUD appears next to the text caret, or at the top or bottom of the screen, or not at all (Settings › General › HUD); nine colour themes, optional Liquid Glass background.
- Sound cues on recording start, stop and cancel (Settings › General › Output).
- Nowhere to paste: when no text field has focus, the text stays in the HUD card with Copy and Close (Esc closes). Markdown results and assistant answers open in a separate window, rendered, with a Source view to edit.
- History (opt-in, Settings › General): “History…” in the menu bar opens a floating window with the last 10 dictations and commands. Clicking a card copies it to the clipboard; ▶ pastes it into the active editor.
- Launch at login (Settings › General › Startup) via the system Login Items mechanism.
Speech engines, silence and hallucination filter
- Models download on first use: Whisper via WhisperKit, NVIDIA Parakeet v3 and Parakeet Ultra via FluidAudio, Apple's built-in recognizers, and the optional Gemini API (default model
gemini-2.5-flash; API key from aistudio.google.com in Settings › General; the vocabulary list is passed in the instructions). - Filter silence and hallucinations (on by default, Settings › General): before transcription the Silero voice detector (FluidAudio) cuts out everything that is not speech, and a recording without speech is dropped; afterwards phrases models invent on silence (“Dziękuję.”, “Thank you.”, “Napisy stworzone przez społeczność Amara.org”…) are removed, the ambiguous ones only where the audio under them is not speech.
Automatic updates
Off by default, Settings › General › Updates. Once a day the app looks for a newer GitHub release; when one is found and no dictation has run for a minute, it downloads the zip, checks the bundle identifier, version and signature, replaces itself in place and restarts. “Check for Updates…” in the menu bar checks right away.
Build & run
Scripts/make-app.sh # release build → build/SimpleWhisper.app, then opens it
Scripts/make-app.sh debug # faster debug buildThe app icon is generated by Scripts/make-icon.py (Pillow) into Resources/AppIcon.icns.
Requires Xcode 26 / Swift 6.2. No Xcode project; it is a Swift Package plus a script that wraps the binary in an .app (needed for the permission prompts).
Permissions
System Settings › Privacy & Security:
| Permission | Why |
|---|---|
| Microphone | recording |
| Accessibility | paste via ⌘V, caret position for the HUD, global key listener |
| Input Monitoring | Globe/fn and Esc key detection |
Also set System Settings › Keyboard › “Press Globe key to” → Do Nothing, otherwise macOS opens the emoji picker or system dictation.
Test from the command line
The same binary can run the whole pipeline on an audio file (no microphone or permissions needed):
.build/release/SimpleWhisper --transcribe test.aiff --engine whisperSmall
.build/release/SimpleWhisper --transcribe test.aiff --engine parakeetV3 --language pl # or pl,en,de / any
.build/release/SimpleWhisper --transcribe test.aiff --prompt "Clean up" --clipboard "some code"
.build/release/SimpleWhisper --transcribe test.aiff --live # split at pauses like Live typing
.build/release/SimpleWhisper --transcribe test.aiff --no-filter # without the silence/hallucination filter
.build/release/SimpleWhisper --apple-locales # which locales Apple Speech supports/installed
.build/release/SimpleWhisper --ax-probe # what the focused element in the front app looks like
.build/release/SimpleWhisper --hud-demo # cycles the HUD through its animations
.build/release/SimpleWhisper --settings-demo # opens the Settings view in a plain window--hud-demo also takes SW_HUD_LIVE=1 (Live typing editor), SW_HUD_RESULT="text" (result card), SW_HUD_THEME, SW_HUD_PLACEMENT and SW_HUD_SOLID=1. The last recording is kept as ~/Library/Application Support/SimpleWhisper/last-recording.wav for --transcribe.
Generate test audio with the system voices: say -v Zosia "Dzisiaj testuję enova365" -o test.aiff.
Notes and limitations
claude -pruns with no tools by default ({tools}in the command template becomes--tools ""): fast, and the model cannot read local files. Settings › AI › Command mode can allow WebFetch/WebSearch.claude -pneeds several seconds (about 4 s with--model haiku --setting-sources "", about 10 s with defaults). The HUD shows an elapsed-seconds counter while it runs.- Do not add
--bareto theclaude -pcommand: it skips OAuth login and fails with “Not logged in”. - Apple Intelligence (Foundation Models) does not support Polish text, so prompts default to Claude Code CLI. Apple Intelligence works for English-only text.
- Apple Speech: English uses the new
SpeechTranscriber; Polish is not supported by it, so the olderDictationTranscriberis used (lower quality). Whisper Large v3 Turbo or Parakeet v3 give much better Polish. - Language auto-detection is restricted to the selected languages (default English) so short phrases are not mistaken for a similar language. Apple Speech supports only some languages (see
--apple-locales); unsupported ones are skipped. - Whisper decodes in the detected dominant language; English identifiers inside a Polish sentence are kept.
- Data lives in
~/Library/Application Support/SimpleWhisper/(prompts, vocabulary, macros, history as JSON) andUserDefaults. Whisper models are cached in~/Documents/huggingface/models/argmaxinc/whisperkit-coreml/, Parakeet in~/Library/Application Support/FluidAudio/. - HUD placement: native text views report the caret rectangle via Accessibility. Electron apps (VS Code, Slack, Teams) report a bogus 0×0 caret, so the frame of the focused text field is used instead. Diagnostics go to
~/Library/Application Support/SimpleWhisper/debug.log. - The app is signed ad-hoc with a designated requirement based only on the bundle identifier, so Accessibility / Input Monitoring grants survive rebuilds. If they ever get stuck, run
tccutil reset Accessibility pl.wojas.SimpleWhisper(andListenEvent,PostEvent,Microphone).
Project layout
Sources/SimpleWhisper/
App/ entry point, menu bar UI, DictationController (pipeline), Updater, CLI and demo modes
Hotkey/ CGEventTap listener for the dictation key and Esc, permissions helpers
Audio/ AVAudioEngine recorder → 16 kHz mono Float32, pause detection for Live typing
Engines/ SpeechEngine protocol, WhisperKit / Parakeet / Apple / Gemini engines, VAD, hallucination filter
Vocabulary/ custom terms + alias post-processing
Macros/ voice macros and two-stage expansion
AI/ prompt model and providers (CLI, HTTP APIs, Apple Foundation Models)
Output/ paster, selection reader, focus detection, result window
HUD/ floating capsule near the caret, Live typing editor, themes
History/ history window
Settings/ SwiftUI settings window
MIT, see LICENSE. SimpleWhisper is completely free.






