← Back to context

Comment by albert_e

5 hours ago

What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.

What do streaming implementations do when a bigger context reveals a different interpretation/parse? When I use Whisper in the terminal, I can see it going back and correcting itself. Are corrections off the table for a true streaming transcription?

My mind is boggled by how many implementations miss this.

Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!

This is the main reason I lean on Deepgram over local services.

  • You can absolutely do high quality, low latency, even multilingual local streaming nowadays. As the commenter above says, Nemotron 3.5 Streaming is awesome. We make heavy use of it in our transcription app.