minimax/speech-02-hd
Text-to-Audio (T2A) that offers voice synthesis, emotional expression, and multilingual capabilities. Optimized for high-fidelity applications like voiceovers and audiobooks.
Capabilities
Cost
Community model (estimated from hardware time)
Input Parameters
| Name | Type | Description | Default | Constraints |
|---|---|---|---|---|
text* | string | Text to narrate (max 10,000 characters). Use markers like <#0.5#> to insert pauses in seconds. | — | — |
audio_format | string | File format for the generated audio. Choose mp3 for general use, wav/flac for lossless, or pcm for raw bytes. | "mp3" | mp3wavflacpcm |
bitrate | integer | MP3 bitrate in bits per second. Only used when audio_format is mp3. | 128000 | 3200064000128000256000 |
channel | string | mono for 1 channel (default), stereo for 2 channels. | "mono" | monostereo |
emotion | string | Desired delivery style. Use auto to let MiniMax choose, or pick a specific emotion. | "auto" | autohappysadangryfearfuldisgustedsurprisedcalmfluentneutral |
english_normalization | boolean | Improve number/date reading for English text (adds a small amount of latency). | false | — |
language_boost | string | Optional language hint. Choose Automatic to let MiniMax detect the language, or pick a specific locale. | "None" | NoneAutomaticChineseChinese,YueCantoneseEnglishArabicRussianSpanishFrenchPortugueseGermanTurkishDutchUkrainianVietnameseIndonesianJapaneseItalianKoreanThaiPolishRomanianGreekCzechFinnishHindiBulgarianDanishHebrewMalayPersianSlovakSwedishCroatianFilipinoHungarianNorwegianSlovenianCatalanNynorskTamilAfrikaans |
pitch | integer | Semitone offset applied to the voice (−12 to +12). | 0 | min: -12, max: 12 |
sample_rate | integer | Audio sample rate in Hz. | 32000 | 80001600022050240003200044100 |
speed | number | Speech speed multiplier (0.5–2.0). Lower is slower, higher is faster. | 1 | min: 0.5, max: 2 |
subtitle_enable | boolean | Return MiniMax subtitle metadata with sentence timestamps (non-streaming only). | false | — |
voice_id | string | Voice to synthesize. Pick any MiniMax system voice (e.g. English_Wiselady, English_Deep-VoicedGentleman) or a voice_id returned by https://replicate.com/minimax/voice-cloning. See the full list of voices in the README. | "English_Wiselady" | — |
volume | number | Relative loudness. 1.0 is default MiniMax gain. Range 0–10. | 1 | min: 0, max: 10 |
textrequiredstringText to narrate (max 10,000 characters). Use markers like <#0.5#> to insert pauses in seconds.
audio_formatstringFile format for the generated audio. Choose mp3 for general use, wav/flac for lossless, or pcm for raw bytes.
"mp3"bitrateintegerMP3 bitrate in bits per second. Only used when audio_format is mp3.
128000channelstringmono for 1 channel (default), stereo for 2 channels.
"mono"emotionstringDesired delivery style. Use auto to let MiniMax choose, or pick a specific emotion.
"auto"english_normalizationbooleanImprove number/date reading for English text (adds a small amount of latency).
falselanguage_booststringOptional language hint. Choose Automatic to let MiniMax detect the language, or pick a specific locale.
"None"pitchintegerSemitone offset applied to the voice (−12 to +12).
0min: -12, max: 12sample_rateintegerAudio sample rate in Hz.
32000speednumberSpeech speed multiplier (0.5–2.0). Lower is slower, higher is faster.
1min: 0.5, max: 2subtitle_enablebooleanReturn MiniMax subtitle metadata with sentence timestamps (non-streaming only).
falsevoice_idstringVoice to synthesize. Pick any MiniMax system voice (e.g. English_Wiselady, English_Deep-VoicedGentleman) or a voice_id returned by https://replicate.com/minimax/voice-cloning. See the full list of voices in the README.
"English_Wiselady"volumenumberRelative loudness. 1.0 is default MiniMax gain. Range 0–10.
1min: 0, max: 10b2c687e53557Updated: 7/25/20262.4M runs
cinemasetfree.com