Elevenlabs - Multilingual V2

ElevenLabs

ElevenLabs

ElevenLabs on FuseAITools—five audio workflows: Multilingual v2 and Turbo 2.5 TTS, Speech-to-Text, Sound Effect v2, and AI Audio Isolation. Voice synthesis, transcription, SFX, and stem separation in the browser.

Multilingual v2
Turbo 2.5
Speech-to-Text
Sound Effect v2
AI Audio Isolation

Text-to-Speech Multilingual v2 Configuration

Required. Select a voice for speech generation. Click the speaker icon to play a sample.
Supports multilingual text, max 5000 characters
0/5000
Variable (0)Stable (1)
Low (0)High (1)
Natural (0)Dramatic (1)
Slow (0.7)Fast (1.2)
Values below 1.0 slow down speech, above 1.0 speed it up. Extreme values may affect quality.
Whether to return timestamps for each word in the generated speech
Optional. Can be used to improve speech continuity when concatenating multiple generations. Max 5000 characters.
0/5000
Optional. Can be used to improve speech continuity when concatenating multiple generations. Max 5000 characters.
0/5000
Optional. Language code (ISO 639-1) to enforce a language for the model.
Select audio output format and quality

Generation Result

No speech generated yet

Enter text and click "Generate Speech" to start synthesis
🎤Text-to-Speech: Supports multiple languages and voice styles, adjustable stability, similarity and style parameters
📝Speech-to-Text: High-precision speech recognition with speaker identification and audio event marking
🎵Sound Effect Generation: AI-driven sound effect generation with loop playback and duration control
✂️AI Audio Isolation: Intelligently isolate vocals and background music
🎙️ ElevenLabs · Voice & Audio · Five Workflows

🌍 ElevenLabs Multilingual v2 TTS

Convert text to natural speech with a required voice selection and up to 5000 characters of input. Tune stability, similarity boost, style, and speed (0.7–1.2). Optional previous/next context text, language code, and word timestamps for subtitle workflows.

What is ElevenLabs Multilingual v2 TTS?

ElevenLabs Multilingual v2 on FuseAITools (elevenlabs_text_to_speech_multilingual) synthesizes speech from text. Required: voice (from the curated voice list) and text (max 5000 chars). Optional: stability/similarity/style/speed sliders, timestamps, previous/next context (5000 chars each), ISO language_code, and output format (MP3 128–320 kbps or PCM). Priced per 1K characters.

🎙️ ElevenLabs on FuseAITools

ElevenLabs on FuseAITools covers professional voice and audio in the browser—natural text-to-speech (Multilingual v2 or fast Turbo 2.5), accurate speech-to-text with optional diarization, AI sound effects, and audio isolation for stems. Pick a voice, paste or upload audio, and download results from history. Credits appear on the Generate button before you submit; new users receive 20 free credits on sign-up.

✨ ElevenLabs Core Features

Two TTS Models

Multilingual v2 for premium prosody; Turbo 2.5 for faster, lower-latency speech.

Speech-to-Text

Upload up to 200MB—auto language, diarization, event tags, and clickable word timelines.

Voice Fine-Tuning

Stability, similarity, style, and speed sliders plus optional context text for seamless multi-clip narration.

Cloud on FuseAITools

Generate, transcribe, and isolate in the browser—credits shown before submit; no local GPU required.

🎯 Built for These Scenarios

Video & podcast narrationDubbing & accessibilityMeeting transcriptsGame & film SFXVocal stem extraction

📊 ElevenLabs Workflow Quick Guide

WorkflowInputBest for
Multilingual v2 TTSVoice + text (≤5000 chars)High-quality narration, dubbing, audiobooks
Turbo 2.5 TTSVoice + text (≤5000 chars)Low-latency voice for assistants and batch runs
Speech-to-TextUploaded audio (≤200MB)Transcripts, subtitles, meeting notes
Sound Effect v2Text description (≤5000 chars)Game, video, and UI sound design
AI Audio IsolationUploaded audio (≤10MB)Vocal/instrument stems and remix prep

❓ FAQ (ElevenLabs)

Browse the Voice dropdown and click the speaker icon to preview samples. Each voice includes a short description—pick one that matches your language, tone, and use case (narration, character, news, etc.). Voice selection is required before generating.

⚙️ ElevenLabs Technical Specs

Parameters below match the FuseAITools ElevenLabs form and API (elevenlabs_* model keys).

WorkflowmodelKeyRequiredPricing unitKey controls
Multilingual v2 TTSelevenlabs_text_to_speech_multilingualVoice + textcredits / 1K charsStability, similarity, style, speed (0.7–1.2); optional timestamps, context text, language code; MP3/PCM output
Turbo 2.5 TTSelevenlabs_text_to_speech_turboVoice + textcredits / 1K charsSame voice controls as Multilingual v2—optimized for faster generation
Speech-to-Textelevenlabs_speech_to_textUploaded audio URLcredits / minLanguage (auto or ISO code); speaker diarization; audio event tagging; word timeline in results
Sound Effect v2elevenlabs_sound_effectSound descriptioncredits / minDuration 0.5–22s; loop toggle; intensity (prompt influence); MP3/PCM output
AI Audio Isolationelevenlabs_audio_isolationUploaded audio URLcredits / minUpload MP3/WAV/M4A (≤10MB)—isolates vocals or instruments from mixed audio

TTS text and context fields: max 5000 characters each. STT uploads: max 200MB. Isolation uploads: max 10MB. Output formats include MP3 (128–320 kbps) and PCM (16–44.1 kHz).

ElevenLabs — Five Audio Workflows

Pick the mode that matches your starting material:

💳 New users get 20 free credits on sign-up. View pricing for subscription discounts and credit top-ups.
🔗 ElevenLabs pipeline tip Narrate with Multilingual v2 or Turbo 2.5 , transcribe meetings via Speech-to-Text , design SFX with Sound Effect v2 , and split stems with AI Audio Isolation . For full music tracks, pair with Suno Music Generation .