Supported Models
Below are the MiniMax speech models and their key features.Supported Languages
MiniMax TTS models provide strong cross-lingual capabilities, supporting 40 widely used global languages. Our goal is to break language barriers and build truly universal AI models.Text Normalization
Control how written text is converted into spoken text before speech generation. Text normalization converts written forms into forms that are more natural to speak. For example, numbers, dates, times, currencies, phone numbers, and other structured expressions may need to be expanded or interpreted based on their context. Iftext_normalization_mode is omitted, the voice_setting.text_normalization setting applies (disabled by default).
You control the normalization strategy with the text_normalization_mode field:
basic— fast, rule-based text normalization with low latency.quality— higher-quality text normalization using a hybrid LLM + rule-based approach, with additional latency.
Basic mode
basic is designed for low-latency text normalization.
It uses a rule-based normalization system to handle common written expressions before speech generation.
For most TTS requests, basic provides a good balance of normalization quality, stability, and latency.
Quality mode
quality is designed for scenarios where text normalization quality is more important than latency.
It uses a hybrid LLM + rule-based approach. Based on the input text, the system automatically routes the request to either the rule-based branch or the LLM branch.
Compared with basic, this mode provides better normalization for more complex or context-dependent inputs.
Currently,
quality takes effect only for non-streaming HTTP requests (stream=false) with speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, and speech-2.6-turbo. For streaming requests (including WebSocket) or other models, quality is processed as basic and no error is returned.Latency
quality introduces additional preprocessing latency:
- approximately 30 ms when the rule-based branch is selected
- approximately 100 ms when the LLM branch is triggered
Which mode should I use?
For most applications, we recommend usingbasic.
Use quality when normalization accuracy is more important and the additional preprocessing latency is acceptable.
Endpoints
All three synchronous TTS APIs support the following domains. For US West,api-uw.minimax.io is recommended, with reduced Time to First Audio (TTFA).
Streaming Request Example
This guide demonstrates streaming playback of synthesized audio while saving the full audio file. ⚠️ Note: To play audio streams in real-time, install MPV player first. Also, ensure your API Key is set in the environment variableMINIMAX_API_KEY.
Request example:
Recommended Reading
Synchronous Text to Speech (HTTP)
Submit the full text in a single HTTP request and get the synthesized audio back, with optional streaming output.
Synchronous Text to Speech (WebSocket)
Send text sentence by sentence over a WebSocket connection and receive streamed audio, suited for low-latency real-time playback.
Bidirectional Streaming Text to Speech (WebSocket)
Stream text over a WebSocket connection at any granularity (even character by character); the server buffers it into sentences before synthesis, ideal for piping LLM streaming output into speech.
Pricing
Detailed information on model pricing and API packages.
Rate Limits
Rate limits are restrictions that our API imposes on the number of times a user or client can access our services within a specified period of time.