GizAI
ACE-Step 1.5
Modalities
text →
Music model for structured songs with style tags, lyrics, BPM, key, language, and long durations.
OmniVoice
Massively multilingual zero-shot TTS with Voice Clone and Voice Design modes, prompt preprocessing, speed, duration, and diffusion controls.
MiniMax
MiniMax Music 1.5
Song model for natural vocals and rich arrangements with English or Chinese structured lyrics.
Stability AI
Stable Audio 3 Medium
Fast high-quality music and sound generation with audio-to-audio editing, inpainting, and continuation.
Inflect Micro v2
Instant CPU-efficient English narration with one consistent synthetic male voice.
Inworld TTS-1.5 Mini
Low-latency expressive TTS for responsive characters, assistants, and real-time narration.
xAI Text-to-Speech
Expressive multilingual TTS with named voices and inline tags for pauses, emphasis, and style.
OpenAI TTS (Text to Speech)
Reliable OpenAI text-to-speech for clear narration, dialogue, and assistant voice output.
Fish Audio
Fish Audio S2.1 Pro
text → audio
Capabilities
Text To Audio
Formats
MP3 · WAV · FLAC · OGG
Released
Jun 1, 2026
Fish Audio S2.1 Pro is a flagship text-to-speech model built for highly expressive, low-latency speech generation. It supports natural-language bracke
xAI
Mar 16, 2026
xAI Text-to-Speech converts text into natural-sounding spoken audio with a single API call. It offers more than two dozen expressive voices, inline sp
Inworld AI
Inworld Realtime TTS-2
May 5, 2026
Inworld Realtime TTS-2 is a conversational text-to-speech model built for realtime voice interaction rather than static narration. It supports free-fo
Google
Gemini 3.1 Flash TTS
Apr 15, 2026
Gemini 3.1 Flash TTS is a text-to-speech model for expressive spoken audio generation from text. It supports granular control over delivery through au
MiniMax Music 2.6
Apr 10, 2026
MiniMax Music 2.6 is MiniMax’s latest music generation model for full vocal songs and instrumentals from text prompts. It supports natural-language pr
MiniMax Music Cover
text / audio → audio
Audio To Audio
MiniMax Music Cover is MiniMax’s song-to-song transformation model for reimagining an existing track in a new style. It preserves the original vocal m
AI model
ACE-Step v1.5 XL Base
Duration
30–300 sec
ACE-Step v1.5 XL Base is the 4B DiT variant of ACE-Step 1.5 for high-quality music generation and editing. It supports text-to-music, cover generation
ACE-Step v1.5 XL SFT
ACE-Step v1.5 XL SFT is the supervised fine-tuned 4B DiT variant in the ACE-Step 1.5 XL line. It is positioned as the highest-quality XL option, combi
ACE-Step v1.5 XL Turbo
ACE-Step v1.5 XL Turbo is the accelerated 4B DiT variant of ACE-Step 1.5. It is optimized for faster music generation with 8-step distilled inference
MiniMax Speech 2.8
Jan 29, 2026
MiniMax Speech 2.8 is an advanced text-to-speech model that turns text into natural, expressive audio in multiple languages. It delivers broadcast-rea
Inworld TTS-1.5 Max
Jan 21, 2026
Inworld TTS-1.5 Max is a high-fidelity text-to-speech model engineered for expressive voice synthesis with rich prosody, nuanced emotional range, and
Inworld TTS-1.5 Mini is a lightweight text-to-speech model designed for real-time voice experiences with ultra-low latency and efficient performance.
Alibaba
Qwen3-TTS 1.7B Base
Jan 1, 2026
Qwen3-TTS 1.7B Base is the foundation text-to-speech model from Alibaba's Qwen3-TTS family. It generates human-like speech across 10+ languages includ
Qwen3-TTS 1.7B CustomVoice
Qwen3-TTS 1.7B CustomVoice is a text-to-speech model from Alibaba that offers nine premium preset timbres across various combinations of gender, age,
Qwen3-TTS 1.7B VoiceDesign
Qwen3-TTS 1.7B VoiceDesign is a text-to-speech model from Alibaba that creates custom voices from natural language descriptions specifying emotion, to
ACE-Step v1.5 Base
ACE-Step v1.5 Base is an open-source music generation foundation model built on a hybrid LLM planner and Diffusion Transformer architecture. It genera
ACE-Step v1.5 Turbo
ACE-Step v1.5 Turbo is a speed-optimized variant of the ACE-Step v1.5 music generation model. It delivers faster inference with fewer denoising steps
Dia2 2B
Nov 19, 2025
Dia2 2B is a 2 billion parameter streaming text-to-speech model from Nari Labs designed for real-time conversational AI. It begins generating audio im
Mirelo
Mirelo SFX 1.5
text / video → audio
1–10 sec
MP3 · WAV · FLAC · OGG · MP4 · MOV · WEBM
Mirelo SFX 1.5 converts video into synchronized sound effects. It targets higher audio fidelity and wider scene coverage. It helps developers add cont
ByteDance
Seed Audio 1.0
text / image / audio → audio
Text To Audio · Audio To Audio
Jun 4, 2024
Seed Audio 1.0 is a ByteDance speech generation model built for high-quality, highly controllable audio output. Public technical materials describe th
fish-audio
S1
Jul 30, 2026
S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetic
S2 Pro
S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language
microsoft
MAI-Voice-2-Flash
→
Jul 23, 2026
MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other
alibaba
Qwen-Audio-3.0-TTS Flash
Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesize
Qwen-Audio-3.0-TTS Plus
Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.
deepgram
Aura-2
Jul 17, 2026
Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multipl
minimax
Speech 2.8 HD
Jul 16, 2026
MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary
Speech 2.8 Turbo
MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitr
MAI-Voice-2
Jun 3, 2026
MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, educatio
canopylabs
Orpheus 3B
Apr 24, 2026
Orpheus 3B is an English text-to-speech model from Canopy Labs, fine-tuned for natural prosody and expressive delivery. It offers 7 preset voices and
sesame
CSM 1B
CSM 1B is a conversational speech model from Sesame. It accepts text input and produces English speech output, with voice options spanning conversatio
hexgrad
Kokoro 82M
Kokoro 82M is a lightweight, open-weight text-to-speech model from hexgrad. It converts text to speech across 8 languages (American and British Englis
mistralai
Voxtral Mini TTS
Apr 19, 2026
Voxtral Mini TTS is Mistral's text-to-speech model featuring zero-shot voice cloning and multilingual support. It converts text input into natural-sou
openai
TTS-1
Nov 6, 2023
TTS is a model that converts text to natural sounding spoken text.
TTS-1 HD
TTS is a model that converts text to natural sounding spoken text. The tts-1-hd model is optimized for high quality text-to-speech use cases.
Select a "Live now" model to generate for free, or connect an API for premium models.