Inworld收购Ultravox整合语音与AI代理平台
Inworld把语音合成和AI代理平台整合在一起,开发者现在能构建更自然的语音交互体验,TTS-2模型还支持情感化语音指令。
Inworld AI收购了Ultravox,现在同时拥有语音合成和AI代理运行平台。Ultravox提供实时语音代理构建服务,支持理解、推理、工具调用和回复功能。收购后Ultravox客户立即升级到Inworld的Realtime TTS-2语音模型,该模型在Artificial Analysis排行榜名列前茅,支持200多种语言,响应时间低于100毫秒。
. @Inworld has acquired @ultravox_dot_ai, so one company now owns both the voice an AI agent speaks with and the platform the agent runs on.
Ultravox is a platform developers use to build real-time voice agents, which are AIs people talk to out loud, like a support line, a tutor, or a companion app.
An agent like that has to do several things at once: understand what the person said and what they meant, reason about it, call whatever tools it needs to actually get the task done, and reply. Ultravox handles all of that.
It also handles the part that makes these agents feel robotic when it goes wrong, which is timing. The agent has to tell the difference between someone pausing to think and someone actually being done, and if the user cuts in mid-sentence it has to stop talking and listen instead of plowing on.
Inworld is an AI research lab and inference provider, meaning it trains its own speech models and runs LLMs for other companies through APIs used by some of the largest consumer AI apps.
Ultravox customers get the first benefit right away, because every built-in Inworld voice on the platform is now upgraded to Realtime TTS-2, Inworld's text-to-speech model.
TTS-2 is top-ranked on the Artificial Analysis leaderboard, an independent ranking where listeners compare voices blind and vote for the better one.
It supports 200+ languages and starts speaking in under 100 milliseconds, meaning roughly a tenth of a second passes between the text being ready and the first sound coming out.
It also takes natural-language direction, so developers can write plain instructions inside the script, like "yell at the top of your lungs with anger" or "slow down and speak with a warm whisper" and the model changes its pacing, tone, and emotion to match.
The longer-term goal is speech-to-speech: one system that listens, reasons, and speaks, rather than three separate steps handing off to each other.