I wrote a blog post for anyone building AI models, rethinking human-computer interaction, or just looking to build truly great speech products:
• Why TTS is finished as a model paradigm—and what’s next with multimodal speech synthesis.
• How model evaluations must evolve to focus on communication, not just naturalness.
• Why product success hinges on understanding speech’s purpose: meaningful user intent.
https://www.papercup.com/blog/speech-ai