What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.
How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)
This definitely seems lighter and faster. How does accuracy compare?
I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.
For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.
Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.
Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.
Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.
FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.
Futo only produces source-available proprietary software. They most certainly are not Open Source, though they unfortunately lied about this a lot before they got called out enough times.
Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.
Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.
What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.
How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)
This definitely seems lighter and faster. How does accuracy compare?
I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...
For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.
There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.
i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes
Unrelated: I love your username.
This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.
I initially read this as whistle to text, which would be way cooler.
A good project for training an llm
https://en.wikipedia.org/wiki/Silbo_Gomero
I've met folks that descend from the Zapotec in southern Mexico, and they still use whistling language to talk to each other across mountain valleys.
Just like Marvel's Yondu!
Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.
Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
RuntimeError: audio limit is 30 s
Did you record over 30 seconds?
Those are insane benchmarks at this size. Well done!
love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.
https://github.com/kouhxp/yapsnap/
Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...
Spanish from where? Here (Castilian Spanish) it seemed to work fine.
Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.
Mexican and venezuelan aren't detected correctly
eager to see if working in android phones
FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.
https://futo.tech/
I'd classify it as "source-available" given the noncommercial clause in the license.
https://github.com/futo-org/android-keyboard/blob/master/LIC...
Futo only produces source-available proprietary software. They most certainly are not Open Source, though they unfortunately lied about this a lot before they got called out enough times.
https://github.com/futo-org/voice-input/blob/master/LICENSE....
I've used FUDO keyboard for a long time, but I never tried using the voice input, so I'm testing it now. Let's see how well it transcribes everything.
...Okay that was pretty good.
Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.
Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.
Can we have more languages?