Free, open-source voice dictation and AI assistant for Windows, macOS, and Linux. Offline Whisper speech-to-text: everything Wispr Flow and Superwhisper do, plus screen vision. Press a key and speak into any app. Say "Hey Flow" and it writes the reply from what's on your screen.
# Add to your Claude Code skills
git clone https://github.com/AbhishekBarali/SpeakoFlowLast scanned: 8/8/2026
{
"issues": [
{
"type": "npm-audit",
"message": "@huggingface/transformers: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "kokoro-js: Vulnerability found",
"severity": "high"
},
{
"type": "npm-audit",
"message": "sharp: sharp inherited vulnerabilities in libvips: CVE-2026-33327, CVE-2026-33328, CVE-2026-35590, CVE-2026-35591",
"severity": "high"
},
{
"file": "AGENTS.md",
"line": 79,
"type": "secret-exfiltration",
"message": "Instruction appears to send credentials/secrets to an external endpoint",
"severity": "medium"
}
],
"status": "WARNING",
"scannedAt": "2026-08-08T04:56:06.722Z",
"npmAuditRan": true,
"pipAuditRan": true,
"promptInjectionRan": true
}SpeakoFlow is an open-source ai agents skill for AI coding assistants such as Claude Code, Codex CLI, and ChatGPT, built by AbhishekBarali. Free, open-source voice dictation and AI assistant for Windows, macOS, and Linux. Offline Whisper speech-to-text: everything Wispr Flow and Superwhisper do, plus screen vision. Press a key and speak into any app. Say "Hey Flow" and it writes the reply from what's on your screen. It has 106 GitHub stars.
SpeakoFlow returned warnings in SkillsLLM's automated security scan. It has no critical vulnerabilities, but review the flagged issues in the Security Report section before adding it to your workflow.
Clone the repository with "git clone https://github.com/AbhishekBarali/SpeakoFlow" and add it to your Claude Code skills directory (see the Installation section above).
SpeakoFlow is primarily written in Rust. It is open-source under AbhishekBarali on GitHub, so you can review or fork the full source.
Yes. SkillsLLM lists many other AI Agents skills you can browse and compare side by side. Open the AI Agents category from the badge at the top of this page, or use the Related Skills and comparison links further down to weigh SpeakoFlow against similar tools.
No comments yet. Be the first to share your thoughts!
Requires a passing catalog security scan. Resolve the flagged issues and resubmit to enable featuring.
A free, local voice assistant for your desktop. Dictation, writing, and an AI assistant, all by voice.
All releases · Website · Documentation
Get told when there's a new version: click Watch → Custom → Releases at the top of this page.
SpeakoFlow turns your voice into text, right where you're working. Press a hotkey and talk, and your words are typed into whatever app you're using. Say "Hey Flow" to turn what you say into a finished reply or email, or open a floating assistant panel to chat by voice and get answers read back to you.
Speech-to-text runs locally on your machine, so your voice never leaves your device. The AI assistant runs on any model you choose, from a fully offline built-in model to your own local server or a cloud provider with your own key. You decide how much stays on your machine.
I built it while studying alone for exams. I was paying for dictation software that stopped at typing: it could hear me, but it couldn't help me.
Most dictation tools stop at typing. Wispr Flow, Superwhisper, and Handy all turn speech into text well. None of them can look at what you are working on and write the reply for you.
SpeakoFlow does both. It's also the only one of the four that's free, open source, and runs on all three desktop platforms.
If all you need is dictation, Handy is a solid choice. If you want your computer to answer you, keep reading. See also: the best free and open-source Wispr Flow alternatives.
Begin a dictation with "Hey Flow" and SpeakoFlow acts on what you said instead of transcribing it. Describe the email, reply, or draft you want and it writes the finished text and pastes it where your cursor is. The trigger phrase is renameable, and it works in any app that accepts text. This is the part plain dictation tools don't do.
Ask a question about whatever you're looking at and the assistant answers with that context: the error in your terminal, the contract in your browser, the chart in your spreadsheet. Combined with Generate with Flow, it can write a reply based on what's on screen rather than on what you dictate. It only captures when you ask it to, the capture goes only to the model provider you chose, and only a small thumbnail is kept locally.
Press a hotkey and talk. Words type into any app, live as you speak or all at once when you stop. Transcription runs on your GPU or CPU with whisper.cpp or Parakeet, fully offline.
A floating always-on-top chat you open with a hotkey. Ask by voice or text, get streaming answers, and have them read back aloud. Collapses to a pill when you don't need it.
Speak another language and get clean English, on your device, with a Whisper model. No cloud round-trip.
Remove filler words and fix grammar in a tone you choose: Professional, Friendly, Concise, or your own custom instruction.
Optional web search so the assistant can look things up for current, factual answers. Profiles switch it between personas, each with its own voice and reply length. Personal memory is on-device and optional, so it learns how you like to work. It's off until you turn it on, and you can edit or erase it at any time.
Everything lives in Settings, and every hotkey is rebindable.
Full documentation for each: Generate with Flow, screen vision, dictation, the assistant panel, languages and translation, AI cleanup, web search, profiles, and memory.
| Action | Windows | macOS | Linux |
|---|---|---|---|
| Dictate | Left Ctrl + Left Super |
Option + Space |
Ctrl + Space |
| Ask the assistant | Left Ctrl + Left Alt |
Option + Ctrl + Space |
Ctrl + Alt + Space |
Hold the shortcut to talk and release to type it out, or switch Recording behavior to Tap in Settings so one press starts and the next press stops. Tap is the hands-free option. The choice applies to every recording shortcut, and all shortcuts are rebindable.
Every shortcut and its default, on all three platforms: Keyboard shortcuts.
Download the latest build for Windows, macOS, or Linux from the Releases page. A short setup wizard helps you pick a transcription model and, optionally, a local model for the assistant.
Download the .exe installer and run it. Windows may show a SmartScreen notice
because the installer isn't signed by a known publisher yet. Choose More
info → Run anyway.
yay -S speakoflow-bin
# or
paru -S speakoflow-bin
.deb
and install it. This registers the app icon and menu entry properly, which the
AppImage can't do on its own:
sudo apt install ./SpeakoFlow_*_amd64.deb
The .deb is built on Ubuntu 24.04, so it needs that era of glibc. On an
older release, use the AppImage instead.chmod +x), and run it. Note that an AppImage
doesn't integrate with your desktop by itself, so it won't show an icon in your
file manager or app menu; tools like Gear Lever or AppImageLauncher add that if
you want it.The AppImage and .deb are both built for x86_64 and ARM64. There's no .rpm
yet, because the packaging doesn't bundle the speech engine correctly, and
shipping one that installs but can't transcribe would be worse than not shipping
it.
Download the .dmg and drag SpeakoFlow into Applications. macOS then needs
Microphone and Accessibility permissions (System Settings → Privacy &
Security) so SpeakoFlow can hear you and type into other apps.
Because the app isn't Apple-signed yet, macOS blocks the first launch and needs one Terminal command to clear it. Full explanation below, or in the install docs.