crazygringo

I think we're about to see a dictation revolution.

I've been coding with Claude, and I pretty much do it exclusively by voice now, simply because it's so much faster. And I kind of feel like Scotty from Star Trek IV, when he holds up a Mac mouse to try to talk into it.

I've been using VoiceInk [1] which looks like it's basically the same as this, but has been around for longer.

What has really made it work for me is using a Bluetooth media control like [2], where I use Karabiner Elements to remap its play/pause button to the dictation keyboard shortcut. I map rewind to option-backspace to delete the last word, fast-forward to shift-enter to insert a newline, a lower button to enter to submit my prompt, and the volume up/down buttons to scroll up/down. I've also just ordered to Xiaomi Bluetooth remote [3] that includes a microphone itself to see if I can get it to work by speaking directly into it and using it as the microphone -- there are a couple of open source projects to turn it into a Mac microphone directly. Since I'd like to be able to talk more quietly instead of into my Mac.

But what I'm REALLY waiting for is the ability to use one button for dictation, and a second button for issuing commands for a local LLM to interpret. So I hold down the voice button which transcribes "I think we need to catch that" and then the second button and go "change catch to cache, like c-a-c-h-e". Or hold down the second button and go "switch to VS code". I don't want something as finicky as macOS Voice Control, I want a local LLM I can speak naturally to.

I can very much see a future where I spend the majority of my "work" time using a Apple TV-type remote.

[1] https://github.com/Beingpax/VoiceInk

[2] https://www.amazon.com/Satechi-Bluetooth-Multimedia-Remote-C...

[3] https://www.notebookcheck.net/Xiaomi-Bluetooth-Remote-2-Pro-...

show comments
Zizizizz

This reminds me very much in terms of features as https://github.com/cjpais/Handy https://handy.computer , what's the use case for this over that?

show comments
rcarmo

As much as I like seeing this, the space is really crowded, and I think the real value is not in single dictation for input, but in two other things: diarization (essential for call recordings, and a staple in video calling services, but hardly touched for in-room meetings and brainstorming) and the next step, which is cleaning it all up for actually useful notes. Not just the transcript, but better variations on meeting notes that you can customize, link to previous existing information, etc.

show comments
Royce-CMR

What’s really shocking is this and many of the competitors are all using OpenAI whisper - which came out in 2022. (Once upon a time, OpenAI really did open research)

It’s surprising in 2026 this is still the default “best” all rounder. Kudos to the early OpenAI team and the open community efforts / improved models since then.

show comments
scosman

If folks like local-AI dictation, but want it optimized for meetings recordings (like Granola/Otter/Notion) check out Biscotti. Free, local, separates voices, voice identification, AI summaries, etc.

https://github.com/scosman/Biscotti

show comments
NJL3000

This came up a month ago - I forked from something now gone - https://github.com/NickJLange/parrot

And if MacOS 26 does it right - all of this is no longer necessary for us to vibe fork/ code on weekends :-)

P.S. Handy looks slick. May be able to ditch mine

show comments
handfuloflight

What about adding a wake word detection so I can just say "Hey wakeWord" to activate the STT? Then maybe detect silence to finish. Completely hands free?

xnx

Seems everyone is cobbling these together. Here's a local-only one from Google: https://apps.apple.com/us/app/google-ai-edge-eloquent/id6756...

maxpert

Been using Handy I am pretty happy but good to see more options, specially OSS.

show comments
sajithdilshan

Even though this looks really good I don’t see myself using this.

Mostly because I could use this only when I’m working from home, and even then after a few minutes I’ll get tired of speaking to myself and I actually type way faster than I speak.

GodelNumbering

Just tried it out as I was looking for something that does this exact thing well (macbook). So far I dictated about 5 different things in different contexts, it worked without a single error! Good choice to pick parakeet v3.

Nice work, thanks for building it!

show comments
jasonjmcghee

The native / built-in API SpeechAnalyzer (introduced last generation on iOS and MacOS) works quite well.

Are folks finding it lacking / needing alternatives?

show comments
frenchie4111

This is awesome. I really want to get something like this integrated directly into my agent IDE (https://getness.dev).

The main thing I need is some way of auto-sending a "end of message" button (like a newline would do) at the end of the message so that it auto-sends the chat.

Not sure the exact right way to implement that, but it seems like something that could be general purpose. Maybe a setting per-active app?

show comments
jakobov

This is cool. I also built a dictation app that has advantage of using three state-of-the-art models in parallel. Which increases accuracy and reduces p95 significantly. You can check it out at https://zwhispr.com/

show comments
gekoxyz

I expected dictation to be much more useful on the computer, but since I can type decently fast I find myself using it far less than I expected.

show comments
Benard-dev

How are you handling local transcription? We opted for an offline-first for FetchMark archiver for same privacy reason and it doubled our retention.

Congrats on shipping!

show comments
kartik017

just in case here is the repo: https://github.com/opennookorg/betterwispr

do star it on gh if you liked it!

epaga

I'm a bit confused, isn't "Wispr" a trademarked term? Why would you pick a name that could easily be shut down?

show comments
Na6z

This is pretty neat, I built something similar a couple months ago. I'd say try to support Windows or Linux next.

show comments
aucisson_masque

How does it compare to handy ?

show comments
starik36

This is my entry into the world of vibe coded Wispr Flow clones. Runs locally on the Whisper AI speech recognition model file - the small variety - just under 200mb.

I thought the small model won't be any good, but it's actually great and barely causes the GPU any usage.

Like other people in this thread mentioned, I switched pretty much exclusively to working with voice. Whether it's responding to text or talking to Claude or writing an email.

https://imgur.com/YJyz05r

honkycat

I've been trying to figure out how to code and build while on the bike /treadmill at the gym.

Want a remote terminal into my Mac that I can voice control.

show comments
ProofHouse

To be honest, I used Wspr so much I only recently started realizing just HOW bad it gets trying to 'improve' your diction, especially coding. I have seen totally opperate directives. I had no idea it was potentially occuring, but it gets past a point where it is WAY to confident in guessing what you meant, sometimes to the extent of changing DON'T to DO, for example.

show comments
stevyhacker

[dead]

stevyhacker

[dead]

467593457091

[dead]

kartik017

Hey HN,

I’m building BetterWispr, a free and open-source voice dictation app for macOS.

I wanted to make voice typing more accessible without requiring a monthly subscription or sending every recording to a cloud server.

You can hold Option + Space, speak naturally, and have the transcription inserted directly into whatever app you’re using.

A few things I’ve been working on:

- Local speech recognition using Whisper, Parakeet, or Apple Speech - Automatic cleanup of filler words and repeated phrases - Custom vocabulary for names and technical terms - Different writing styles depending on the app - Optional meeting transcription and summaries

The project is open source under Apache 2.0.

I’m still improving the experience, especially around transcription accuracy, speed, and reliability across different Macs.

I’d love feedback from people who regularly use voice dictation.

What would make you switch from your current dictation tool to an open-source alternative?

Website: https://betterwispr.com/

show comments