Train your own AI voice with Ultimate TTS. This complete Windows tutorial covers voice cloning, automatic take selection, multilingual speech and OmniVoice full fine-tuning. The app also supports RunPod, SimplePod, Massed Compute and local Linux.
Links:
- Full YouTube Tutorial link : [** https://www.youtube.com/watch?v=ivppvgjqAtg **]
- Download Ultimate TTS and read the source post: [** https://www.patreon.com/SECourses/posts/ultimate-text-to-139297407 **]
- Windows Requirements tutorial: [** https://youtu.be/DrhUHnYfwC0 **]
Video Chapters:
- 00:00:00 My trained AI voice
- 00:00:38 600+ supported languages
- 00:01:19 Windows and cloud options
- 00:01:46 Your first voice clone
- 00:02:10 Generate and listen
- 00:03:18 Download ZIP and guides
- 00:03:46 Windows Requirements
- 00:04:28 Choose an install path
- 00:04:40 Extract the package
- 00:04:54 Run the installer
- 00:05:23 PyTorch, CUDA and wheels
- 00:05:54 Start the application
- 00:06:19 Models, voices, GPU presets
- 00:07:04 First load vs warm runs
- 00:07:46 Players and saved outputs
- 00:08:34 IndexTTS English example
- 00:09:02 AuK English example
- 00:09:52 Reference library and files
- 00:10:48 Select a reference range
- 00:11:38 Microphone recording
- 00:11:53 Reference text and transcription
- 00:13:19 Clone and Auto voice
- 00:14:38 Voice design tags
- 00:15:26 Language selection
- 00:15:40 Preview and Smart sentences
- 00:16:32 Section-length study
- 00:17:01 Sentence and section pauses
- 00:18:08 Pause tags
- 00:19:04 Check and set pronunciation
- 00:19:38 Save and reuse dictionary
- 00:21:10 Expressive tags
- 00:21:44 Keep which take
- 00:21:58 Fewest word errors
- 00:22:12 Rank takes and set checks
- 00:22:42 Ranking and fallback result
- 00:23:49 Whole-text candidates
- 00:24:36 60-line selection study
- 00:25:05 139-line Auto voice study
- 00:25:35 First vs selected take
- 00:26:42 Matching-language references
- 00:27:26 Arabic voice example
- 00:28:24 Chinese voice example
- 00:29:04 French voice example
- 00:29:46 German voice example
- 00:30:40 Italian voice example
- 00:31:22 Japanese voice example
- 00:32:03 Korean voice example
- 00:32:58 Polish voice example
- 00:33:38 Spanish voice example
- 00:33:49 Cross-language cloning
- 00:35:48 Speaking-rate comparison
- 00:36:49 Target duration modes
- 00:37:39 Best quality: 32 steps
- 00:38:25 Fast: 16 steps
- 00:38:54 Seeds and advanced sampling
- 00:39:44 WAV, MP3 and filenames
- 00:40:01 Reference and subtitle export
- 00:40:16 Inspect the exported files
- 00:41:10 Recent outputs as references
- 00:41:52 Audio tuning comparison
- 00:43:09 Tuning and edge silence
- 00:43:30 SRT import and cue timing
- 00:44:22 Create a narrated MP4
- 00:45:08 Batch inputs and filenames
- 00:45:43 Shared and per-file voices
- 00:46:07 Batch execution and progress
- 00:46:47 Inspect batch results
- 00:48:00 Save a personal preset
- 00:48:18 Load settings and reference
- 00:49:26 Recent settings and theme
- 00:49:45 Select GPU and memory tier
- 00:50:02 BF16 and ConvRot INT8
- 00:50:16 Measured memory use
- 00:50:30 Measured synthesis time
- 00:50:44 Compare both variants
- 00:52:25 Section batch size
- 00:52:39 Load models and find files
- 00:53:03 VRAM benchmark and tier cap
- 00:53:56 Train a new OmniVoice
- 00:54:15 Source recordings and text
- 00:54:44 Reserve evaluation sessions
- 00:55:00 Dataset folder and scan
- 00:55:37 Transcript and alignment
- 00:55:57 Segment duration
- 00:56:13 Padding and caption cleanup
- 00:56:28 Quality filters
- 00:56:47 Prepare and inspect clips
- 00:58:02 Speaker and transcript audit
- 00:58:50 Review flagged clips
- 00:59:21 Correct words and boundaries
- 01:00:20 Build the OmniVoice cache
- 01:00:47 Dataset and training name
- 01:01:06 Full fine-tuning
- 01:01:33 Full, LoRA and DoRA
- 01:01:50 Held-out voice comparison
- 01:02:05 Epochs and Training plan
- 01:02:22 Token budget and accumulation
- 01:02:39 Learning rate and memory
- 01:03:14 Checkpoints and train state
- 01:03:30 Continue vs weights-only
- 01:03:47 Evaluation and final test
- 01:04:18 Fluency analysis and filters
- 01:05:16 Start training
- 01:05:49 Speed, loss and validation
- 01:06:18 Hear a training sample
- 01:06:45 Stop and resume controls
- 01:07:05 800-update completion
- 01:07:39 Checkpoint analysis
- 01:07:55 Automatic speech comparison
- 01:08:15 Reference audition
- 01:08:37 Speaking-rate calibration
- 01:08:53 Save an INT8 voice copy
- 01:09:46 Load the trained preset
- 01:10:14 Generate held-out speech
- 01:11:12 Base vs trained voice
- 01:12:15 Configure Checkpoint Grid
- 01:13:24 Hear the six grid cells
- 01:14:34 Trained Auto voice
- 01:15:22 Trained INT8 voice
- 01:16:13 Updates, backup and help
- 01:17:01 Your next voice
This video is for anyone who wants a complete voice workflow, from a short reference to a reusable trained voice. Start with the included reference and save your first working preset.
For support, questions, feature requests, installation issues and updates, use the download post, pinned comment, Patreon, Discord and video comments. Thank you for watching.
Ultimate Text To Speech Generator With Voice Cloning turns your scripts, subtitles and whole chapters into natural speech on your own PC, with three speech models in one app: IndexTTS 2.5, OmniVoice and AuK. Clone any voice from a short clip, design a new voice from a few words, edit recordings, and train your own voice from your videos. Below you will see every feature with real screenshots from the app.
OmniVoice DoRA (better version of LoRA) trainings is yielding literally better than ElevenLabs with our special pipeline + Keep which take inference system.
Why Ultimate Text To Speech Generator
- Three speech models in one app: IndexTTS 2.5, OmniVoice and AuK. Switch them in the header; each keeps its own settings.
- Clone any voice in seconds: drop an audio or video file and type your text. On an RTX 5090, IndexTTS 2.5 made 31 seconds of speech in 18 seconds from a 14.6 second part of a video.
- Take quality: the app renders several takes of every section, Whisper checks them, and the take that sounds most like the voice without a word error is kept. In the developer's test on 60 lines, word errors fell from 0.65 % to 0.20 % with OmniVoice.
- New voices without a recording: describe a voice in one sentence with AuK, or pick voice tags with OmniVoice.
- AuK Audio Editing: 23 editing tasks: replace, add or remove words, change speed, pitch, volume or emotion, clean up noise and separate voices. Replacing two words in a 10 second clip took 1.7 seconds.
- Train your own voice for any model: from your videos and subtitles to a finished voice with a ready-to-use preset. Four recordings became 205 clean training clips in 50 seconds.
- Made for long narration: Smart sentences, exact pause tags, a live section preview and a pronunciation dictionary.
- Subtitles to speech and MP4: 10 subtitle formats, speech timed to every cue, and an MP4 with your image in the same run.
- Batch generation: whole folders of scripts and subtitles in one click.
- Ready for your GPU: 7 GPU presets from 6 GB to 32 GB, BF16 or INT8 ConvRot models, block swap and a built-in VRAM benchmark.
- 1-click installers: Windows, RunPod, SimplePod, Massed Compute and Linux, with PyTorch 2.14.1 and CUDA 13. The models download automatically, and every model weight file is checked with SHA256.
- Frequent updates: four releases between 2 and 4 October 2026, with a built-in Help tab and changelog.
Here is the app right after a voice clone. The voice came from a 14.6 second part of one of our tutorial videos, and 31 seconds of speech were ready in 18 seconds.
1: choose IndexTTS 2.5, OmniVoice or AuK in the header. 2: the preset for your GPU is selected on the first start. 3: any audio or video file can be the voice. 4: type your script. 5: the result plays right away, with the elapsed time, speed and VRAM next to it.
1-Click Installation
What You Download
You get one small zip file with the installers for every platform.
Extract it into a short folder path without spaces, double-click Windows_Install_or_Update.bat, then start the app with Windows_Start_App.bat. Run the same installer again at any time to update.
Latest PyTorch and CUDA 13
The installer makes its own Python 3.12 virtual environment, so your other apps stay untouched. It installs PyTorch 2.14.1 with CUDA 13, Transformers, Triton for Windows and every other package with uv, a very fast package installer.
PyTorch, torchvision and torchaudio always come from the same CUDA 13 build. Our fresh install took 5 minutes 39 seconds on our PC, models included.
Automatic Model Downloads
At the end, the installer downloads the IndexTTS 2.5 models and their helper models by itself.
Large files download over 16 connections, and every model weight file is checked with SHA256, so a broken download is caught right away. Running the installer again skips the large files you already have.
Start the App
Double-click Windows_Start_App.bat. On the first start, the app reads your GPU memory and selects the matching GPU preset.
The app was ready in 9.5 seconds on the first start and in 4 to 6 seconds later. OmniVoice, AuK and the built-in Whisper download automatically the first time you use them.
Cloud GPUs: RunPod, SimplePod and Massed Compute
You can also run the app on a cloud GPU. The cloud installers set up everything with one command and download the models for you. A Gradio share link lets you use the app from any device.
- SimplePod: register here and use this template.
- RunPod: register here and use this template.
- Massed Compute: register here and use our coupon SECourses.
The step-by-step commands are in the instruction files inside the zip.
Requirements
On Windows you need Python 3.12, Git, FFmpeg, CUDA 13, cuDNN 9.17 and Visual Studio with the C++ tools. This tutorial shows every step, and the same setup runs all our AI apps.
Three Speech Models in One App
Every model has its own strengths, and you switch between them with one click in the header.
IndexTTS 2.5 gives you emotion control and reads five languages. OmniVoice clones voices, designs voices from tags and speaks many languages. AuK clones voices, designs voices from a plain sentence and edits recordings in its own tab. Each model keeps its own settings, and every preset stores all three.
GPU Presets for 6 GB to 32 GB Cards
Seven read-only presets fit the whole app to your card: generation, dataset preparation and training together. On the first start, the app selects the preset of your GPU.
Where to find it: the Universal preset list sits at the top of the page, above every tab.
Open the Universal preset list to see every preset:
Change anything you like, type a name and click Save to keep your own preset. It stores every setting of every tab and of all three models. Every training adds a ready-to-use preset of its new voice to this list. The Help tab shows what each GPU preset uses:
A smaller card pays with speed first: helper models move to the CPU, and the 6 GB preset streams 22 of the 24 GPT blocks from RAM.
Voice Cloning From Any Audio or Video
A clean clip of a few seconds is enough to clone a voice. Audio and video files both work, and the preview plays exactly what the app will use.
Where to find it: open the Voice Generation tab. Reference Voice is the left column.
Here is the Reference Voice column up close:
Drop a file, record from your microphone, or load any file on your PC by its path. Time ranges such as 0.2:14.8 keep only the clean part of a long recording, and several ranges are joined in order.
Text, Smart Sentences and Pauses
Long text is split into sections before the voice reads it. Smart sentences packs whole sentences into each section.
Where to find it: in the Voice Generation tab, Text & Timing is the middle column.
Here are the text controls:
Tags such as [pause:800ms] insert an exact silence. With a trained voice, the line length and the pauses between sentences come from the speaker's own recordings. Below the text, the live section preview shows every section before you generate:
You see the tokens and words of each section and every pause, so you can fix the text before you spend any GPU time.
Pronunciation Check and Dictionary
Names and technical words are the hardest part of any narration. The pronunciation check finds the words the model may not know and suggests a reading for each.
Where to find it: in the Voice Generation tab, click Pronunciation check & dictionary under the live section preview.
Click Check unknown words:
Readings come from the CMU dictionary, acronym and CamelCase splitting, and letter rules. Save the good ones to your dictionary, and the app applies them in Voice Generation, Batch Generation and the live preview, for IndexTTS 2.5 and OmniVoice.
Take Quality: The Best Take of Every Section
Take quality picks the best take of every section. It renders several takes, lets the built-in Whisper check every word, and keeps the take that sounds most like the voice without a word error. It works for all three models, and the GPU presets switch it on.
Where to find it: in the Voice Generation tab, Take quality is in the right column, under the live log.
Here is Take quality after an OmniVoice clone of the same video clip:
OmniVoice rendered 10 takes of our 32 second script in 26.5 seconds, and the app kept the most similar take with 0 word errors. In the developer's test on 60 lines of voice cloning, word errors fell from 0.65 % to 0.20 % with OmniVoice and from 0.91 % to 0.35 % with IndexTTS 2.5.
OmniVoice: Clone, Design or Auto Voice
OmniVoice clones a reference voice, designs a new voice from tags, or picks a voice by itself.
Where to find it: choose OmniVoice in the header. Its section opens in the Voice Generation tab, under Voice LoRA / DoRA.
Here is the OmniVoice section with a designed voice:
We picked four tags: female, young adult, moderate pitch and british accent. The new voice spoke 6.1 seconds of speech in 1.2 seconds, with no reference recording at all.
AuK: Describe a Voice in Words
AuK clones voices, designs a voice from one plain sentence, and lets a trained AuK voice speak without a reference.
Where to find it: choose AuK in the header. Its section opens in the Voice Generation tab, under Voice LoRA / DoRA.
Here is the AuK section with a described voice:
"A calm middle-aged man with a deep, warm voice, speaking slowly and clearly, like a documentary narrator." AuK made 6.6 seconds of speech from that sentence in 1.5 seconds. AuK speaks English and Chinese.
Emotion Control
IndexTTS 2.5 separates the voice from the emotion, and the app gives you four ways to set the emotion.
Where to find it: with IndexTTS 2.5 selected, scroll down in the Voice Generation tab and click Emotion Control. The Open / close all sections button at the top opens it too.
Here is Emotion Control with the emotion sliders selected:
Keep the speaker's own tone, copy the delivery of an emotion clip, mix joy, anger, sadness, fear, disgust, depression, surprise and calm, or describe the emotion in words. Emotion weight sets how strong it is.
Timing, Output Formats and Word Timestamps
Make the speech fit your video, save it in the format you need, and get subtitles of every word.
Where to find it: in the Voice Generation tab, click Segmentation & Timing, Output and Execution.
The three sections opened:
Ask for an exact length, such as 25 seconds, and Natural mode paces the speech to fit. Every run saves a WAV file, one tick adds an MP3 up to 320k, and another adds SRT and VTT subtitles with the timing of every word. Isolated subprocess mode frees all VRAM after each job.
Subtitles to Timed Speech and MP4
Turn any subtitle file into speech that follows its timing, ready for dubbing, translations and slide videos.
Where to find it: in the Voice Generation tab, the Captions and Still image for MP4 boxes are under the text controls.
Drop your subtitle file and tick Use caption cue timing:
SRT, VTT, SBV, ASS/SSA, SUB, LRC, TTML/DFXP, SAMI, JSON and TSV all work, and the app recognizes the format from the file itself. Add an image, and the same run also makes an MP4:
Our 6-cue subtitle file became exactly 30.00 seconds of speech, with every cue starting on time.
AuK Audio Editing
AuK edits real recordings: change the words, the delivery or the sound, and separate voices.
Where to find it: choose AuK in the header, then click the AuK Audio Editing tab.
Here is a word replacement:
We replaced "quick test" with "short demo" in a 9.9 second recording. The edit took 1.7 seconds, and Whisper heard the new words in the result. The 23 tasks also insert or remove words, change speed, pitch, volume and emotion, turn speech into a whisper, add a laugh, clean up noise and reverberation, and keep one speaker or the voices from music.
Batch Generation
Generate many scripts or subtitle files in one run.
Where to find it: click the Batch Generation tab.
Drop your files, paste your texts or point the app at a folder, then click Generate batch:
Our three stories became 116 seconds of speech in 65 seconds with OmniVoice, every section picked from 10 takes. Use one voice for all files, or a matching audio file next to each script for a different voice per file. Continue after item errors keeps the batch going.
Train Your Own Voice
A trained voice follows the speaker much more closely than a single reference clip: the pace, the pauses and the way each word is said. You can train IndexTTS 2.5, OmniVoice and AuK, and the app takes you from videos and subtitles to a finished voice.
Prepare the Dataset
Point the app at your videos or audio files with their subtitles. It cuts them into clean clips of whole sentences.
Where to find it: click the LoRA Dataset Preparation tab.
Here is a real run on four of our tutorial recordings:
Your subtitles give the words and Whisper gives the exact word timing. The four recordings, 46.8 minutes in total, became 205 clean clips of 40.6 minutes in 50 seconds. Every clip can be checked before training:
The clips are 4 to 16 seconds long, and every clip keeps its text.
Voice and Transcript Audit
The audit checks every clip against a clean recording of the speaker and against its text, and builds a separate training dataset from the clips that pass.
Where to find it: in the LoRA Dataset Preparation tab, click Voice and transcript audit at the bottom.
Here is the audit of our 205 clips:
It checked 205 clips in 1 minute 14 seconds and kept 194. A second Whisper model rescued 7 good clips that the first one misheard. A whole recording can be held out, so the app tests the voice on speech it never trained on.
Start the Training
Choose your dataset, give the voice a name and click Start training. The defaults are the settings we measured on real voices.
Where to find it: click the Voice Training tab.
Here is the tab with our audited dataset:
The GPU VRAM preset fits the training to your card. DoRA and LoRA train small adapters; full fine-tuning trains the speech model's own weights on 16 GB cards and larger (AuK 32 GB). The fluency filter can train on fluent clips only. During training the dashboard updates live:
Our IndexTTS 2.5 voice trained at about 5 to 6 steps per second with 2.3 GB of VRAM. The app checks the voice on held-out clips every epoch and keeps the best checkpoint.
Automatic Checks After Training
When training ends, the app keeps working for you. It renders the same sentences with the base model and with the checkpoints, then compares them with the speaker's real recordings of those sentences. It also calibrates the speaking rate, trains a voice decoder adapter, tunes the decoding settings and auditions reference clips. Here is the comparison for our voice:
Likeness to the speaker's real recordings rose from 0.810 with the base model to 0.867 for the chosen checkpoint with its voice decoder adapter, and the word error rate fell from 2.4 % to 2.3 %. The app picked that checkpoint by itself and set it as the voice's default.
A Ready-to-Use Preset
Every training ends with a ready-to-use preset of the new voice: our training saved my_voice_Takes_5 in the Universal preset list you saw above. It loads the chosen checkpoint, the best reference clip from the audition, the calibrated speaking rate of 0.86, the decoding settings from the sweep, and Take quality with up to 5 takes per section.
Use Your Trained Voice
Select your voice in Voice Generation, and the app loads everything it learned.
Where to find it: in the Voice Generation tab, the Voice LoRA / DoRA panel is right under the three columns. You can also load the voice's ready-made preset from the Universal preset list.
Here is the panel after loading the preset of our new voice:
The panel shows the voice's calibrated pace, the line length and the pauses measured from its own training clips, and the decoding its sweep adopted, and applies them to your text. With this preset our voice read the 88-word script as 35.2 seconds of speech in 28 seconds, and the built-in Whisper found 0 word errors in every section.
Compare Checkpoints by Ear
The Checkpoint Grid tab renders the same texts with every checkpoint you choose, so you can listen to them side by side.
Where to find it: click the Checkpoint Grid tab.
Here is a grid of our voice with the base model and three checkpoints:
Every cell uses the same texts, reference and seed, so you hear only the difference between the checkpoints. Every grid is saved and opens again at any time, and Calibrate speaking rate from this grid measures the voice's real pace.
Models and Performance
Runtime
Fit the models to your GPU in one place.
Where to find it: click the Models & Performance tab.
Here are the runtime controls:
Choose a VRAM tier, and the fit check shows the expected peak memory before anything loads. Pick the BF16 model or the INT8 ConvRot model, which saves memory. Unload model / free VRAM clears the GPU in one click.
Block Swap and Model Placement
Small cards can stream parts of the model from system RAM.
Where to find it: in the Models & Performance tab, click Block Swap & Memory and Auxiliary Model Residency.
Here is the 6 GB tier:
The 6 GB tier streams 22 of the 24 GPT blocks from RAM and keeps the reference models on the CPU. Every helper model can stay on the GPU, stay on the CPU or load on demand.
Model Files and VRAM Benchmark
Check your model files and test a GPU tier before a long job.
Where to find it: in the Models & Performance tab, click Model Files & Downloads and VRAM Benchmark.
Both sections after a real test:
The INT8 ConvRot GPT is 1.1 GB instead of 3.1 GB and downloads from inside the app. The VRAM benchmark ran the 8 GB tier with its memory limit in place on our card: it fit with a 3.5 GB peak and made 25.9 seconds of speech in 19.3 seconds.
Built-in Help and Changelog
You never need to leave the app to learn it.
Where to find it: click the Help tab or the Changelog tab.
The Help tab starts with a four-step quick start:
It explains every model, every workflow and every parameter. The Changelog tab lists every release in plain words:
The newest release is always at the top.
Latest Updates
The app gets frequent updates. Versions 7.0 to 8.1 came out between 2 and 4 October 2026:
- 8.1: fresh installs keep PyTorch, torchvision and torchaudio on one CUDA 13 build, the preset after training carries the voice's calibrated speaking rate and decoding, IndexTTS training caches its features by itself, and the chosen speech model stays selected after a reload.
- 8.0: Take quality for all three models, a ready-to-use preset after every training, and OmniVoice as the default model of a new install.
- 7.1: AuK, the third speech model, with voice design from a sentence, trained voices without a reference, and the AuK Audio Editing tab.
- 7.0: OmniVoice, the second speech model, with voice design, BF16 or INT8 ConvRot and training.
Get Ultimate Text To Speech Generator
Download the latest zip file attached to this post, extract it and run the installer for your platform. To update, get the newest zip, overwrite the old files and run Windows_Install_or_Update.bat again.
- Requirements tutorial: Python, Git, FFmpeg, CUDA and C++ tools step by step.
- Support: our Discord channel.
























































Top comments (0)