Performance metrics of Kakao’s omni AI model “Kanana-o.” Photo courtesy of Kakao
Kakao (035720.KS) has upgraded its self-developed omni AI model “Kanana-o” to generate speech that expresses specific speaking styles, intonation and even emotion. The company also improved generation speed and efficiency through a self-developed voice tokenizer.
According to Kakao’s tech blog on the 5th, the Kanana-o development team has advanced its voice generation technology so that users can control speaking style, emotion and even intonation using natural language commands alone. While existing voice AI focused on reading text naturally, the improved Kanana-o controls the voice expression itself according to the user’s commands.
For example, when a user states a desired manner of speaking in natural language, such as “read it very fast,” “read it in a low voice,” “read it in a sad voice,” or “read it in the Gyeongsang dialect,” the AI generates speech that finely reflects not only speed, volume and pitch but also emotion, intonation and intensity.
In addition, the model can perform role-based commands such as “like a sports broadcast,” “like an announcer,” or “like reading a fairy tale,” as well as complex commands containing multiple conditions, such as “lower the tone and read it fast in a sad voice.”
Kanana-o scored 94.50 on the Korean-language “InstructTTSEval” benchmark, which evaluates the ability to follow speaking commands. This surpasses GPT-4o-mini-tts (91.10 points) and is comparable to Google’s Gemini-2.5-flash-preview-tts (95.38 points). In addition, although Kanana-o was trained primarily on Korean data, it also performs the same speaking commands without difficulty in English voice generation.
The application of the self-developed voice tokenizer “LM-SPT” improved not only the quality of voice generation but also its speed and efficiency. LM-SPT is a technology that allows the AI to compress and express voice with fewer tokens, reducing the amount of data the AI must process so it can generate voice more quickly and efficiently.
Kakao plans to continue advancing its Kanana-o voice technology. Through research that processes voice understanding and generation in a single integrated structure, the company plans to advance its voice processing technology and focus on providing a natural, seamless voice experience. It also plans to continue development for even more detailed voice control, including technology that naturally generates non-verbal expressions such as laughter, sighs and exclamations.
Noh Byung-seok, performance leader of Kakao’s Unified Foundation Model, said, “This advancement of Kanana-o voice technology focused on equipping the model with the ability to generate speech as natural as a human, as well as to express the speaking style, emotion and intonation that users want according to natural language commands.” He added, “Going forward, we will apply the Kanana-o model to various services to provide a more natural and convenient AI voice experience.”