The newly introduced models—focused on transcription, voice synthesis, and image creation—are designed to deliver faster performance, multilingual capabilities, and enhanced creative output, with built-in safety features and access through Microsoft’s developer platforms.
Microsoft has introduced a new suite of artificial intelligence models aimed at expanding capabilities across speech recognition, voice generation, and image creation. The models—MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2—are now available through the company’s MAI Playground and Microsoft Foundry platforms.
The announcement underscores Microsoft’s continued push to strengthen its AI ecosystem by offering tools that cater to developers, content creators, and enterprises seeking advanced automation and creative solutions.
Focus on speed, accuracy, and multilingual support
Among the newly launched tools, MAI-Transcribe-1 is designed to convert spoken language into text across 25 widely used global languages. The company claims the model performs effectively even in challenging audio environments, such as recordings with background noise or unclear speech.
Microsoft stated that the transcription model delivers significantly improved speed, operating up to 2.5 times faster than earlier versions. The service is priced at $0.36 per hour, making it accessible for developers building applications that require real-time or large-scale transcription capabilities.
Meanwhile, MAI-Voice-1 focuses on generating natural-sounding speech with the ability to replicate tone and emotional nuance. The model allows users to create customised voices within seconds and is capable of producing up to 60 seconds of audio in just one second. Priced at $22 per one million characters, it is positioned as a suitable solution for podcasts, virtual assistants, and voice-enabled applications.
Enhancing creative workflows with AI imaging
The third model, MAI-Image-2, is targeted at designers, photographers, and digital creators. It is engineered to generate high-quality visuals at a faster rate than its predecessor, while maintaining accuracy in elements such as colour balance, skin tones, and embedded text.
According to Microsoft, the model delivers image outputs at roughly twice the speed of earlier iterations. Pricing for MAI-Image-2 is set at $5 per one million text input tokens and $33 per one million image output tokens, reflecting its use in scalable creative workflows.
The company emphasised that all three models incorporate safety mechanisms and are developed with a focus on responsible AI usage. By integrating these capabilities into its platforms, Microsoft aims to provide users with flexible tools that combine performance with security.
The rollout highlights the growing competition in the AI space, as technology firms continue to introduce specialised models designed to improve productivity and creativity across industries.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.
