Microsoft (MSFT.US) is accelerating its push to reduce dependency on OpenAI by unveiling two proprietary artificial intelligence models specialized in image and voice generation. The newly introduced models — ‘MAI-Image-2.5-Pro’ and ‘MAI-Voice-2-Flash’ under Microsoft’s ‘MAI’ AI brand — are designed to rapidly replace external models by offering performance and cost efficiency optimized for the company’s own services.
Microsoft’s AI division announced on the 23rd (local time) that it is launching a public preview of both models. MAI-Image-2.5-Pro delivers the highest quality among all image generation models Microsoft has released to date, focusing on demanding tasks such as creating hero images for advertising, detailed editing, and rendering text within images. Pricing is set at $5 per million tokens for text input, $8 per million tokens for image input, and $106 per million tokens for image output.
MAI-Voice-2-Flash targets services that require rapid response and large-scale processing, such as call centers and voice agents. Compared to the existing MAI-Voice-2, processing speed has doubled while pricing has been reduced by 32% to $15 per million characters.
This model launch goes beyond a mere technology demonstration. Microsoft emphasized that it is already achieving meaningful results by deploying its own models across several core products. MAI-Image-2.5 serves as the default image generation model for Bing Image Creator and has also been integrated into image generation and editing features in PowerPoint and OneDrive. According to the company, image transformation costs in PowerPoint have been reduced by up to 84% compared to OpenAI’s GPT-Image-2. In OneDrive, the image editing save success rate improved by 26%, response latency (P95) decreased by approximately 25%, and efficiency in medium-load production environments improved by 2.5 times.
Voice model deployment is also becoming more concrete. MAI-Voice-2-Flash has been integrated into Microsoft’s Dynamics 365 Contact Center, serving customers including T-Mobile (TMUS.US) and UK-based low-cost carrier EasyJet. Microsoft reported that GPU costs were reduced by up to 89% in this process. The model has also been integrated into Azure Voice Live, which supports developers in building voice-to-voice interaction agents.
Mustafa Suleyman, CEO of Microsoft AI, explained the rationale behind the in-house model strategy: “Compared to OpenAI models, our own models run faster, cost less, and deliver higher generation quality, which improves user retention rates.” He specifically cited that MAI model operating costs in PowerPoint have dropped by approximately 85% compared to previous solutions.
This move is widely interpreted as Microsoft’s effort to reduce its technical and financial dependence on specific external models while maintaining its partnership with OpenAI. Microsoft currently holds a long-term agreement allowing free use of OpenAI’s large-scale models, but bears the substantial computing infrastructure costs required to run them. The calculation is that expanding the use of more efficient in-house models can significantly reduce infrastructure expenses.
Indeed, Microsoft appears to be accelerating the transition to proprietary models beyond image and voice, extending into text generation. Market observers note that Microsoft has begun gradually replacing OpenAI and Anthropic text generation models with its own MAI models across various Office product suites. Microsoft has previously stated that its MAI models match the performance of OpenAI’s latest GPT-5.6 model in common tasks within applications like Excel, while offering superior cost efficiency.
“MAI models are now deployed across more than half of Microsoft’s products, and we are conducting tests across the entire product portfolio,” Suleyman said. “This is a very strong start, and we will continue to push forward with full force.” Microsoft also added that its speech recognition model from the same family, ‘MAI-Transcribe-1.5,’ has been applied to the medical voice solution ‘Dragon Copilot,’ supporting 58 languages and reducing transcription error rates by 50% compared to previous models for most languages. With next-generation GB200 computing clusters already operational and signaling continued expansion of the MAI series, Microsoft’s declaration of AI independence is expected to accelerate further.