Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I love the UX of voice mode. I’ve always hated voice models.

Gemini’s being particularly egregious (always ending in some deranged question, can not reliably be prompted away) prompted me to build my own client for my real harness that simply does STT -> model -> TTS (both being independently useful).

I guess I see some value in a model responding quickly and with more nuance, but it’s not much. I can wait for it to finish. I’d much rather have it be actually useful. I’m not looking for a digital friend.

The delegation feature lets me see some value in a voice model for orchestration type features. But in either case, I don’t really like (or understand why others would like) talking to a model with different features and quirks just because I’m using a different medium to communicate over.



I’ve spent my time very similarly working on my own voice stack project, but having also seen how non-developers use AI or experience technology in general, I truly think they are better served with a different UX and product than what we have.

In other words, if you’re building your own voice inference tooling you’re just about the polar opposite user demographic than the one that truly needs and will value this. You’re using voice as a medium of convenience doing what existing models are technically and practically “shaped” to be able to do, knowing how they work well enough that conversation is more like typing/prompting with your voice than a natural interface. I’m guilty of this myself but have you ever even paid for a voice/audio model or hardware?

Compare that to the millions of people with an Alexa device in their home who buy products through it, or who prefer calling support to get a human over poring over technical documentation. They’re actually very close to finally getting a version of “Alexa” that lives up to its promise and I’m happy for them


I have built out something similar that let's me use my phone's hardware buttons to open an input stream with the mic for my Hermes agent over my matrix gateway and then has it play back with a local TTS model on my Pixel 10 Pro.

But the Deepseek v4 flash model I am using through OpenRouter is killing me on latency. Any suggestions to improve that?




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: