Dictate & AI: WhisperShortcut
Speech to Text & Transcribe
Only for Mac
$3.99
Open-source voice-to-text for Mac. Speak anywhere, paste anywhere — bring your own API key or run fully offline with local Whisper. No subscription, no account.
Turn your voice into text anywhere on your Mac. WhisperShortcut lives in your menu bar — press a shortcut, speak, and your words are ready to paste in any app.
Open source. No subscription, no account. Bring your own API key, or run fully offline with local Whisper.
WHAT YOU CAN DO
• Dictate — Speak and get clean text on your clipboard. Use cloud models or local Whisper, fully offline and private.
• Voice editing — Copy any text, then say an instruction like "make this shorter" or "translate to English" to rewrite it.
• Read Aloud — Select text anywhere and hear it in natural AI voices, at the speed you choose.
• AI Chat — A built-in chat with multiple models, screenshots, image and file attachments, web search and slash commands.
• Screenshots — Capture your screen and attach it to a voice prompt or chat.
• Live Meeting — Record and transcribe meetings as they happen, with a local transcript.
• Smart Improvement — The app learns your vocabulary and preferences over time.
PRIVATE BY DESIGN
Local-first, with no backend and no sign-up. Your audio and text go only to the AI provider you choose with your own key — or never leave your Mac with local Whisper.
WORKS WITH YOUR TOOLS
Google Gemini, OpenAI (GPT), and xAI (Grok). Optional Google (Calendar, Tasks, Gmail) and Trello integrations turn chat into a hands-free assistant.
Requires macOS 15.5 or later.
more Connect OpenRouter with one click — sign in instead of creating and pasting an API key. One connection covers both dictation and chat, and you pick models from a live list.
YouTube videos in chat — paste a link and a Gemini model watches the video instead of guessing from search results. Add a timestamp and it analyses a ten-minute window around that moment.
Share Usage Report — a summary of how the app has actually worked for you, containing no transcripts, prompts, replies, or audio. You see the full text before anything is sent.
The transcription model picker now groups models as Direct, Routed and Offline, so it is clear which account each one bills.
Fixes: Dictate Prompt no longer pastes raw JSON into your document when an edit cannot be applied. Gemini 3.1 Pro is no longer offered for dictation, because it fails to answer short recordings. Long, tool-heavy chat turns no longer lose the final answer.
7.98 Aug 3
• Live Meeting: notes now appear in the chat while the meeting runs, and the chat can answer questions about anything said since it started. New: one-tap "Catch me up", quote a note to ask about it, and a shortcut (⌘6) to flag important moments.
• Dictation accuracy: temperature is now configurable and defaults to verbatim — previously it ran at the model default, the likely cause of occasional invented or swapped words. Thinking effort is adjustable too.
• Dictate through any audio-capable model on OpenRouter with a single API key.
7.96 Jul 31
- Bind F1–F20 alone (no modifier) — ideal for a dedicated QMK/VIA dictation key.
- Non-destructive paste: turn on Restore clipboard and whatever you had copied before dictating goes straight back on the clipboard after the text is pasted.
- Copy Last Transcription and Recent Transcriptions (last 5) from the menu bar when paste went to the wrong place.
- Read Aloud starts speaking as soon as the first chunk is ready; markdown, links, and code fences are stripped for clearer speech.
- Better Whisper Glossary homophone handling, fewer false “implausible speech” discards, and Live Meeting no longer flashes Processing Audio on every background chunk.
7.95 Jul 30
- Voice Feedback: correct the app by speaking. Press ⌘5 and say something like "my name is spelled G-o-e-d-d-e" or "stop capitalizing every noun". Your instruction becomes a concrete change to the right part of your context, shown as a diff you approve before anything is applied.
- Grok searches X.com again, and the new /x command narrows it to the accounts you trust: "/x @karpathy @simonw", or "/x off" to search all of X. Set a default under Settings > Chat.
- Chat messages sent while a reply is still generating are now queued instead of replacing it, so no answer gets lost.
- Cancelling a Dictate Prompt now really cancels: a late reply can no longer paste itself into whatever you had focused.
- Lower cost by default: Gemini models now use 3.5 Flash-Lite with tuned thinking levels. Plus fewer wrong glossary learns, and the in-app Chat answers "what can this app do?" from the real feature list and your actual shortcuts.
7.94 Jul 25
- Settings and Chat are now fully navigable with VoiceOver: model tiles, toggles, dropdowns and icon-only buttons all announce what they are and what they control.
- Interface polish across both windows: the Read Aloud voice picker now matches every other model picker, section headers carry icons consistently, and card padding, spacing and button corners are unified.
- Fixed: chat composer text could render nearly black on the dark composer in light appearance, and the Smart Improvement review window clipped its content at its smallest size.
7.92 Jul 24
- Share folders with the chat: it can list, open and search the text files in folders you pick, and it remembers how they are organised. Add one with /folder or by dropping it onto the chat window.
- Change how Dictate Prompt rewrites your text just by asking the chat — no need to open Settings.
- New chat models: GPT-5.6 Sol, Terra and Luna, Grok 4.5, and Claude Fable 5.
- Fixed: correcting a misheard name in the chat now really takes effect in dictation.
7.91 Jul 23
- Two new Gemini models, both cheaper than the ones they replace: Gemini 3.5 Flash-Lite now powers Dictate, and Gemini 3.6 Flash powers Dictate Prompt, Chat and meeting summaries.
- If you never changed your model selection, you move to the new models automatically — a deliberate choice is always kept.
- Fixed: dictation with Gemini 3.1 Pro failed on every attempt due to an invalid API parameter.
- Fixed: the onboarding tour no longer counts as finished when you simply close the window.
7.89 Jul 22
• Fn (Globe) key: tap to toggle dictation, or hold to talk
• API keys no longer vanish from the Keychain — with clear errors if a save fails
• Claude (Anthropic) is now available as a chat provider
• Clearer billing and rate-limit messages for Google, OpenAI, and xAI
• After dictation, an explicit ⌘V cue shows when text is ready to paste
• Safer Google key saving, and clearer help if Accessibility permission looks stuck after switching App Store ↔ GitHub builds
7.88 Jul 19
Google Calendar chat tools now support recurring and all-day events.
• Recurring events: create or update repeating entries like yearly birthdays, weekly standups, and weekday-only series.
• All-day events: birthdays, holidays, and full-day blocks no longer need a start or end time.
• Smarter date handling for single-day all-day events.
7.85 Jul 18
- Dictation is dramatically faster: long recordings are transcribed while you speak, so the transcript arrives moments after you stop. Recordings are also compressed for faster upload.
- Push-to-talk: hold the dictation shortcut to record and release to transcribe — a short tap still toggles recording as before.
- New floating recording indicator with live microphone level and confirm/discard buttons.
- Instant glossary learning: correct a term by typing it in the chat and the dictation glossary learns the spelling immediately.
- Meeting summaries are cheaper and faster.
- "No speech detected" now appears as a brief notice instead of a persistent error popup.
7.81 Jul 9
- Chat replies are easier to read: prose inside code fences now wraps like normal text instead of being cut off, and the composer shows which model is active.
- Fixed a chat freeze that could pin the CPU during streaming replies.
- Better onboarding: the offline Whisper option is always visible, and the final overview shows every shortcut with its real permission status.
- Connect the chat to any OpenAI-compatible endpoint (e.g. self-hosted inference).
- More accurate permission status in Settings and more reliable dictation clipboard handling.
7.78 Jul 5
- Live Meeting uses far fewer API calls — longer transcription intervals and on-demand summaries instead of fixed timers.
- Fixed a chat freeze during long streaming replies.
- More reliable live meeting summaries and clearer AI error messages when limits are hit.
7.76 Jul 4
- Fixed a freeze that could occur during long chat replies — long answers now stay smooth.
- Fixed Live Meeting summaries lagging or piling up during long meetings.
7.75 Jul 3
- New: Run Dictate Prompt fully offline with a local OpenAI-compatible server such as Ollama or LM Studio — just set the endpoint and model in Settings.
- Chat: email questions now get a real answer instead of dead-ending, replies keep a consistent tone, and the sidebar shows Today and Yesterday expanded by default.
- Fix: resolved a Gemini 3 chat error that could break tool calls.
7.74 Jun 29
- Improved welcome tour and onboarding flow, featuring a persistent floating window, a unified permissions hub, and reliable system settings integration.
- Enhanced chat functionality with image downloads, real-time Google integration (Calendar, Tasks, Gmail), and fixes for streaming freezes and Keychain-related stalls.
- Added clearer error handling for non-vision models attempting to read screenshots and overall UI polish.
7.73 Jun 24
- Launch at Login is now off by default — enable it yourself in Settings.
- Auto-paste is now optional; dictation works without extra permissions.
- Clearer setup prompts.
- More reliable AI chat with automatic retries on temporary server errors.
7.68 Jun 18
- More reliable chat: Fewer freezes during long conversations with generated images; smoother streaming and a more stable message list when sending or resizing the window.
- Meeting summaries for every provider: Summaries and titles now work with Gemini, OpenAI, and Grok; missing summaries regenerate when you open the Summary tab.
- Smarter transcription & calendar: Your glossary is applied to instructable speech-to-text models; Google Calendar lookups can include past events with normalized dates.
7.55 Jun 10
- Generate images in AI Chat with Gemini - ask for a picture or let the assistant use the new image tool.
- Chat scrolls more smoothly in long conversations; fixes rare 100% CPU freezes when scrolling.
- Read Aloud shows a speaking indicator in the menu bar; retry your last chat message or copy pasted text with your message.
7.50 Jun 4
- Fixes chat freezing during long responses — especially when switching chats while a response was still generating.
- Up to 10 file attachments per message (previously 5).
7.45 Jun 1
• Read Aloud: more reliable text capture; friendly note when nothing is selected
• Screenshots save to your folder again, with better permission guidance
• Live meeting titles show up in chat sooner
• Fixed chat freezing on replies with web sources
• Restored menu-bar review prompts
7.42 May 31
• Read Aloud: pick a voice per provider (Gemini, OpenAI, Grok) in Settings
• Chat: /think sets reasoning depth per conversation (all providers)
• Read Aloud rewrite sounds more natural with an improved default prompt
7.39 May 30
Read Aloud: Read selected text aloud, optional AI rephrasing for non-readable content, variable speed.
Setup: Record shortcuts by keypress, permissions onboarding, save screenshots and attach in chat.
Chat & Meetings: Sidebar with search and meetings area, faster live transcription, more stable meeting titles.
7.35 May 28
- Paste screenshots and images into Chat; new ⌘3 shortcut captures a region straight to your clipboard.
- Updated chat defaults (GPT-5.5, Grok 4.3); Gemini Pro models use deeper reasoning for stronger answers.
- Smart Improvement runs more reliably in the background with fewer popups; sharper dictation and chat prompts.
7.26 May 25
Dictation: Fixed silent drop to idle (especially offline Whisper); clearer “no speech detected” feedback; more forgiving silence threshold.
Quality & support: Show logs in Settings; better transcription for large non-WAV files; fewer wasted API calls after Dictate Prompt.
Settings: Offline Whisper download confirmation no longer blocks the window; toast stays visible for 10 seconds.
7.20 May 20
This update improves Dictate Prompt with a stronger default model and stricter language and minimal-edit behavior, shows richer attachment details in chat, and refines slash-command recognition (including /copy).
7.16 May 18
Connect OpenRouter with one click — sign in instead of creating and pasting an API key. One connection covers both dictation and chat, and you pick models from a live list.
YouTube videos in chat — paste a link and a Gemini model watches the video instead of guessing from search results. Add a timestamp and it analyses a ten-minute window around that moment.
Share Usage Report — a summary of how the app has actually worked for you, containing no transcripts, prompts, replies, or audio. You see the full text before anything is sent.
The transcription model picker now groups models as Direct, Routed and Offline, so it is clear which account each one bills.
Fixes: Dictate Prompt no longer pastes raw JSON into your document when an edit cannot be applied. Gemini 3.1 Pro is no longer offered for dictation, because it fails to answer short recordings. Long, tool-heavy chat turns no longer lose the final answer.
more Version 7.98 Aug 3
Data Not Collected The developer does not collect any data from this app.