## The first LLM for text-to-speech Prompt to generate AI voices, change emotions, and more Voice Describe the desired AI voice's identity, voice qualities, and more Script To help shape the voice we generate, input something distinctive this AI voice would say #### A text-to-speech system that understands what it's saying Octave (Omni-capable text and voice engine) isn't a traditional TTS model. It’s a voice-based LLM. That means it understands what words mean in context, so it can predict emotions, cadence, and more. <iframe class="absolute inset-0 size-full bg-background-dark" src="https://player.vimeo.com/video/1060564675" title="A text-to-speech system that understands what it's saying" frameborder="0" allow="autoplay; fullscreen; picture-in-picture; clipboard-write" allowfullscreen=""></iframe> #### Create any voice you can imagine with Octave Voice Design Create any AI voice you can imagine, like a "sarcastic medieval peasant," with a brief prompt or evocative script #### Generating the best AI voices has never been easier In a blind comparison study with over 100 human raters, Octave’s outputs were favored over outputs from ElevenLabs Voice Design in terms of audio quality, naturalness, and how well speech generations matched descriptions of the desired voice, across 120 diverse prompts. <iframe class="absolute inset-0 size-full bg-background-dark" src="https://player.vimeo.com/video/1060626591" title="Generating the best AI voices has never been easier" frameborder="0" allow="autoplay; fullscreen; picture-in-picture; clipboard-write" allowfullscreen=""></iframe> #### The first AI voice generator that can take nuanced Acting Instructions As an LLM for voice, Octave can interpret your prompt and adjust its voice accordingly—from “angry” to “just above a whisper” #### Any emotion or speaking style, on command Octave is the first TTS system that can take natural language instructions to change emotional delivery and speaking style. Give directions like "sound sarcastic" or "whisper fearfully." For the [[first time]], creators have total control. <iframe class="absolute inset-0 size-full bg-background-dark" src="https://player.vimeo.com/video/1060626661" title="Any emotion or speaking style, on command" frameborder="0" allow="autoplay; fullscreen; picture-in-picture; clipboard-write" allowfullscreen=""></iframe> #### For creators and developers alike Octave was built to generate the most expressive AI voices for any content: podcasts, voiceovers, audiobooks, and more. With our API, you can bring it to any application. ![TTS Projects](https://directus.hume.ai/assets/1360175d-aa10-44c4-b55c-32b50e0955bd/tts-projects.png?width=1920&height=1920&quality=75&format=webp&fit=inside) #### We research foundation models and how to align them with human well-being Empathic Voice Interface (EVI) ##### Real-time interaction Based on a new voice-to-voice AI model architecture, EVI 2 can converse rapidly and fluently. It understands the user’s tone of voice and generates an appropriate tone of voice automatically. It's capable of emulating a wide range of personalities, accents, and speaking styles. It can replace or integrate with other LLMs. #### Explore EVI 2's capabilities Built for Developers ##### Interact with synthetic voices and personalities Create an interactive personality for your use case with flexible prompting and voice modulation tools. We developed a novel voice modulation approach that allows anyone to adjust EVI 2’s base voices along a number of continuous scales, including femininity, nasality, pitch, and more. Optimized for Human Well-Being ##### Build AI voices people can trust EVI 2 excels at anticipating and adapting to users' preferences, made possible by its special training for emotional intelligence. Its pleasant and fun personality is a result of this deeper alignment with human values. In deploying this technology, we require developers to adhere to the guidelines of The Hume Initiative, a non-profit that sets the first concrete guidelines for empathic AI. ### Developer Resources