Whistle: Speech to Text in 16.9 MB

914 points176 commentsa day ago
skolos

Interesting that this is here. I used whistle (and bunch of other things) to take ownership of my echo show. It now doesn't dial to Amazon at all - it does all processing locally with its own CPU and connects to my homeassistant for home automation. My initial setup involved qwen asr (1.7b model) running on rtx 5080. Compared to that, whistle was really bad (out of 170 messages, qwen recognized correctly 168, whistle - 70), but I adjusted whistle to work like jev - instead of free form transcription it recognizes only select set of templates (I trained tiny network with 10,000 generated utterances to translate whistle final state to probabilities within templates). The precision went up to 164/170 - almost matching qwen. By the way - I'm speaking with heavy accent.

show comments
bambax

I tried the speech by Colonel Jessep at the end of "A Few Good Men"; here's what I got:

> You can handle the truth. We're living a world that has walls and all those walls have to be guarded by men with cannons

This is disturbing. "can handle" instead of "can't handle" may be due to my pronunciation, but "cannons" instead of "guns" is pure hallucination, as well as "all" in "those walls". So maybe don't rely on this for a faithful transcript of what was said.

(For the record, here's the text reference: "You can't handle the truth. We live in a world that has walls, and those walls have to be guarded by men with guns.")

show comments
hn_submit

I find the quality of transcription quite poor with a high error rate. And I'm using an Devine podcast quality microphone.

IMHO the error rate is comparable to Dragon Speech of twenty years ago. It seems to me the speech models don't take into account the context of the words in the sentence. They need to predict what the speaker's going to say so it can compensate for slightly-off pronunciation.

INTPenis

I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

But every single sound he makes with his mouth ends up on the page too.

show comments
albert_e

What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.

show comments
wkcheng

How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)

This definitely seems lighter and faster. How does accuracy compare?

show comments
zimpenfish

Tried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.

show comments
nl

I recently built a 3D printed ESP32 based transcription device[1] that offloads transcription to Parakeet running in the browser connected to the ESP32. I haven't done my own measurements of error rate but it'd be interesting to see how well this runs on the ESP32 itself.

In my case I'm already using one core to run DSP for a beamforming mic array (which works amazingly well for noise cancellation!) so I don't have huge amounts of free processing though.

[1] https://github.com/nlothian/earwing

andai

I made a little gui for it.

(Or rather, ported my little voice typing gui to use Whistle/Needle instead of Parakeet. Only had to change like 2 lines!)

https://gist.github.com/a-n-d-a-i/56ac707eacd8144f83d3651be4...

Seems to do roughly as well as parakeet (both speed and quality) so far.

I use it on Linux. I think it's cross platform but I haven't tested it. Would love to know! Should be trivial (2 mins of Codex) to get it working on other OS if not.

andy_ppp

Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

joewhale

I initially read this as whistle to text, which would be way cooler.

show comments
modinfo

A simple demo for parakeet-redux, which cactus model is based on this tiny model too https://github.com/skorotkiewicz/parakeet-redux-api

flowerlad

Apple needs to incorporate this into iOS ASAP. This works much better than the speech recognition in iOS when you use technical terms. One of the most frustrating parts of iOS is speech-to-text in iMessage. For me no feature is more important in a phone.

Try this example: My website uses ASP.NET technology and I am using .NET 10.0. Works perfectly in Whistle, but not in iOS.

Dave3of5

It's so so 50/50. Tried a few examples and its trying to fit what it thinks it's heard into an actual sentence which means is often inserting or changing words. Don't think I'd used that.

Also who cares about the binary size unless it's like 15 TB or something why not add a bit more MB to the binary size to get a better alog or whatever.

jamienk

Did anyone ever read The First Circle by Solzhenitsyn? It is amazing: in a Stalin gulag in the 1950's USSR, a bunch of Soviet engineers / techies are tasked with creating a speech recognition system. They show fake progress in order to keep from being jailed in worse ways. It is autobiographical, so good, but I read it so long ago and it convinced me how hard speech recognition would be, maybe even impossible!

Highly recommended. I see there is a "new" (2009, I getting old) edition & translation - https://www.amazon.com/First-Circle-Aleksandr-I-Solzhenitsyn...

bwestergard

It does very poorly in detecting Dutch vs. German for this Flemish-accented speaker.

armcat

Those are insane benchmarks at this size. Well done!

amelius

Let's say I want to build a hardware product now, voice-controlled, with voice feedback, so STT, LLM, and TTS. All local. What are the best libraries to do this now, say with 8GB of GPU memory available?

show comments
cedws

Are there such STT models that allow ‘context’ to be provided? I’d love to be able to dictate for programming sessions. The problem is that STT models don’t handle jargon and acronyms well, especially if they’re specific to the project or company. Is there like a type of model you can provide a custom vocabulary to without needing training?

show comments
sleepybrett

This is awesome. What I want now for a little thing I'm working on is 'speech to text with attribution' in an environment with one or more people, assuming no crosstalk I want speech->text and identify the speaker with some kind of internal identification. No prep data at all. 'i heard a person say 'hello there' i haven't heard them before this session, that's person a. 'oh i just heard a new person say 'hello to you', that's person b.

sfpk

If you need speech2text try this, https://ccoreilly.github.io/vosk-browser/

blagui

Error rate is crazy high. And comparing to low models and omit a lot of other bigger open models.

e12e

Hm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish.

Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).

So, reasonable, but limited?

show comments
MisterMunchkin

For something you could plausibly ship inside a webapp, it’s very good. I can definitely think of some cool uses for this.

kamranjon

Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.

TomGarden

The Qwen STT model that's leading the open weight leaderboard right now is excellent. Parakeet v2 is blazing fast and accurate enough on English. I feel like STT gains from this point on will be marginal, especially given that you can do a quick LLM pass afterwards with a small model

RexHuang

Nice work. One gap though: no Chinese in the seven languages — and Chinese is where a lot of the on-device demand actually is. Any plans?

show comments
rafaelm

Huh, this was really confusing. I already had an STT app called Whistle on my Android phone.

flytomoon

Thanks for putting in the work. Only a matter of time before all of this runs locally at very high accuracy. And can handle complex vocabularies.

jakobov

ZWhispr seems to be the best for those who care about accuracy as it uses three SOTA models.

https://zwhispr.com/

show comments
aavisangle

I think it works great on Mac but lags on Windows. The accuracy on Windows is lower than Mac.

jayshah5696

This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.

properbrew

Might look into embedding Whistle into Whistle if it can make it 30x smaller (https://play.google.com/store/apps/details?id=com.blazingban...) - It's a shame there isn't as many languages supported though, I'm surprised at the amount of non-english downloads (I really shouldn't be, of course non-english speakers want dictation) of the app there is.

est

I really hope there's a TTS model in 16.9MB

Kokoro is good, but can't change the pitch easily.

nialv7

Can't even be bothered find a human to write your model announcement article huh.

pzo

tested in polish and unless you speak very loud, clear and slow is not that good, parakeet definitely better.

rshemet

hey, Roman here from Cactus, thank you for the feature!

opening this thread for questions/feedback if you have any

mrkn1

love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.

show comments
cellular

Can it run on any of the cheap uC / esp / stm ?

show comments
mo2art

RuntimeError: audio limit is 30 s

show comments
Centigonal

This is quite good, especially given the size and the fact that it runs on the CPU.

jiajunluo

That's cool, i would want to try it out

saturn8601

Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.

show comments
jakobov

ZWhispr is where its at https://zwhispr.com/

aidotguru

eager to see if working in android phones

show comments
tecleandor

Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...

show comments
timonoko

Google cuboard has improved greatly Lady Lady lately. Now it does almost understand Finnglish. The word "battery" was the worst, now it seems to go all right every time.

villgax

Actually bad WER

snvzz

Ships with a RISC-V build. This is much appreciated!

lab14

Tried it a few times with English, Spanish and French and the quality/accuracy is pretty "meh". If the model doesn't really work, it doesn't matter if it fits in 1MB.

contingencies

For speech to text UX I currently use https://handy.computer/ as it's cross platform and open source. With that I am currently using Parakeet Unified EN 0.6B and finding it excellent. Often I use it to talk to AIs without giving them audio, which works very well. Honestly, I would never go back to typing now. Promised since ~Y2K, the tech is finally here. You really notice it when you wake up at 2AM and don't want to wake people ... it can get really annoying reverting to key-tapping. My long-gnawing fear of losing my hands to RSI is no longer a thing, and I can focus on losing them to another hobby: like sailing or machining! Just bought a band saw...

try-working

I built an STT plugin for DeepSeek Harness that uses this Whistle model as well as a larger one from Desert Ant Labs: https://github.com/try-works/dsh-stt

paaloeye

RIP Wispr Flow

agilek

Can we have more languages?

show comments