DEV Community

Cover image for Better Than ElevenLabs? Ultimate TTS Voice Training & Cloning in 600+ Languages
Furkan Gözükara
Furkan Gözükara

Posted on

Better Than ElevenLabs? Ultimate TTS Voice Training & Cloning in 600+ Languages

Better Than ElevenLabs? Ultimate TTS Voice Training & Cloning in 600+ Languages

Train your own AI voice with Ultimate TTS. This complete Windows tutorial covers voice cloning, automatic take selection, multilingual speech and OmniVoice full fine-tuning. The app also supports RunPod, SimplePod, Massed Compute and local Linux.

Links:

Video Chapters:

  • 00:00:00 My trained AI voice
  • 00:00:38 600+ supported languages
  • 00:01:19 Windows and cloud options
  • 00:01:46 Your first voice clone
  • 00:02:10 Generate and listen
  • 00:03:18 Download ZIP and guides
  • 00:03:46 Windows Requirements
  • 00:04:28 Choose an install path
  • 00:04:40 Extract the package
  • 00:04:54 Run the installer
  • 00:05:23 PyTorch, CUDA and wheels
  • 00:05:54 Start the application
  • 00:06:19 Models, voices, GPU presets
  • 00:07:04 First load vs warm runs
  • 00:07:46 Players and saved outputs
  • 00:08:34 IndexTTS English example
  • 00:09:02 AuK English example
  • 00:09:52 Reference library and files
  • 00:10:48 Select a reference range
  • 00:11:38 Microphone recording
  • 00:11:53 Reference text and transcription
  • 00:13:19 Clone and Auto voice
  • 00:14:38 Voice design tags
  • 00:15:26 Language selection
  • 00:15:40 Preview and Smart sentences
  • 00:16:32 Section-length study
  • 00:17:01 Sentence and section pauses
  • 00:18:08 Pause tags
  • 00:19:04 Check and set pronunciation
  • 00:19:38 Save and reuse dictionary
  • 00:21:10 Expressive tags
  • 00:21:44 Keep which take
  • 00:21:58 Fewest word errors
  • 00:22:12 Rank takes and set checks
  • 00:22:42 Ranking and fallback result
  • 00:23:49 Whole-text candidates
  • 00:24:36 60-line selection study
  • 00:25:05 139-line Auto voice study
  • 00:25:35 First vs selected take
  • 00:26:42 Matching-language references
  • 00:27:26 Arabic voice example
  • 00:28:24 Chinese voice example
  • 00:29:04 French voice example
  • 00:29:46 German voice example
  • 00:30:40 Italian voice example
  • 00:31:22 Japanese voice example
  • 00:32:03 Korean voice example
  • 00:32:58 Polish voice example
  • 00:33:38 Spanish voice example
  • 00:33:49 Cross-language cloning
  • 00:35:48 Speaking-rate comparison
  • 00:36:49 Target duration modes
  • 00:37:39 Best quality: 32 steps
  • 00:38:25 Fast: 16 steps
  • 00:38:54 Seeds and advanced sampling
  • 00:39:44 WAV, MP3 and filenames
  • 00:40:01 Reference and subtitle export
  • 00:40:16 Inspect the exported files
  • 00:41:10 Recent outputs as references
  • 00:41:52 Audio tuning comparison
  • 00:43:09 Tuning and edge silence
  • 00:43:30 SRT import and cue timing
  • 00:44:22 Create a narrated MP4
  • 00:45:08 Batch inputs and filenames
  • 00:45:43 Shared and per-file voices
  • 00:46:07 Batch execution and progress
  • 00:46:47 Inspect batch results
  • 00:48:00 Save a personal preset
  • 00:48:18 Load settings and reference
  • 00:49:26 Recent settings and theme
  • 00:49:45 Select GPU and memory tier
  • 00:50:02 BF16 and ConvRot INT8
  • 00:50:16 Measured memory use
  • 00:50:30 Measured synthesis time
  • 00:50:44 Compare both variants
  • 00:52:25 Section batch size
  • 00:52:39 Load models and find files
  • 00:53:03 VRAM benchmark and tier cap
  • 00:53:56 Train a new OmniVoice
  • 00:54:15 Source recordings and text
  • 00:54:44 Reserve evaluation sessions
  • 00:55:00 Dataset folder and scan
  • 00:55:37 Transcript and alignment
  • 00:55:57 Segment duration
  • 00:56:13 Padding and caption cleanup
  • 00:56:28 Quality filters
  • 00:56:47 Prepare and inspect clips
  • 00:58:02 Speaker and transcript audit
  • 00:58:50 Review flagged clips
  • 00:59:21 Correct words and boundaries
  • 01:00:20 Build the OmniVoice cache
  • 01:00:47 Dataset and training name
  • 01:01:06 Full fine-tuning
  • 01:01:33 Full, LoRA and DoRA
  • 01:01:50 Held-out voice comparison
  • 01:02:05 Epochs and Training plan
  • 01:02:22 Token budget and accumulation
  • 01:02:39 Learning rate and memory
  • 01:03:14 Checkpoints and train state
  • 01:03:30 Continue vs weights-only
  • 01:03:47 Evaluation and final test
  • 01:04:18 Fluency analysis and filters
  • 01:05:16 Start training
  • 01:05:49 Speed, loss and validation
  • 01:06:18 Hear a training sample
  • 01:06:45 Stop and resume controls
  • 01:07:05 800-update completion
  • 01:07:39 Checkpoint analysis
  • 01:07:55 Automatic speech comparison
  • 01:08:15 Reference audition
  • 01:08:37 Speaking-rate calibration
  • 01:08:53 Save an INT8 voice copy
  • 01:09:46 Load the trained preset
  • 01:10:14 Generate held-out speech
  • 01:11:12 Base vs trained voice
  • 01:12:15 Configure Checkpoint Grid
  • 01:13:24 Hear the six grid cells
  • 01:14:34 Trained Auto voice
  • 01:15:22 Trained INT8 voice
  • 01:16:13 Updates, backup and help
  • 01:17:01 Your next voice

This video is for anyone who wants a complete voice workflow, from a short reference to a reusable trained voice. Start with the included reference and save your first working preset.

For support, questions, feature requests, installation issues and updates, use the download post, pinned comment, Patreon, Discord and video comments. Thank you for watching.

Ultimate Text To Speech Generator With Voice Cloning turns your scripts, subtitles and whole chapters into natural speech on your own PC, with three speech models in one app: IndexTTS 2.5, OmniVoice and AuK. Clone any voice from a short clip, design a new voice from a few words, edit recordings, and train your own voice from your videos. Below you will see every feature with real screenshots from the app.

OmniVoice DoRA (better version of LoRA) trainings is yielding literally better than ElevenLabs with our special pipeline + Keep which take inference system.

Why Ultimate Text To Speech Generator

  • Three speech models in one app: IndexTTS 2.5, OmniVoice and AuK. Switch them in the header; each keeps its own settings.
  • Clone any voice in seconds: drop an audio or video file and type your text. On an RTX 5090, IndexTTS 2.5 made 31 seconds of speech in 18 seconds from a 14.6 second part of a video.
  • Take quality: the app renders several takes of every section, Whisper checks them, and the take that sounds most like the voice without a word error is kept. In the developer's test on 60 lines, word errors fell from 0.65 % to 0.20 % with OmniVoice.
  • New voices without a recording: describe a voice in one sentence with AuK, or pick voice tags with OmniVoice.
  • AuK Audio Editing: 23 editing tasks: replace, add or remove words, change speed, pitch, volume or emotion, clean up noise and separate voices. Replacing two words in a 10 second clip took 1.7 seconds.
  • Train your own voice for any model: from your videos and subtitles to a finished voice with a ready-to-use preset. Four recordings became 205 clean training clips in 50 seconds.
  • Made for long narration: Smart sentences, exact pause tags, a live section preview and a pronunciation dictionary.
  • Subtitles to speech and MP4: 10 subtitle formats, speech timed to every cue, and an MP4 with your image in the same run.
  • Batch generation: whole folders of scripts and subtitles in one click.
  • Ready for your GPU: 7 GPU presets from 6 GB to 32 GB, BF16 or INT8 ConvRot models, block swap and a built-in VRAM benchmark.
  • 1-click installers: Windows, RunPod, SimplePod, Massed Compute and Linux, with PyTorch 2.14.1 and CUDA 13. The models download automatically, and every model weight file is checked with SHA256.
  • Frequent updates: four releases between 2 and 4 October 2026, with a built-in Help tab and changelog.

Here is the app right after a voice clone. The voice came from a 14.6 second part of one of our tutorial videos, and 31 seconds of speech were ready in 18 seconds.

Ultimate Text To Speech Generator: Voice Generation tab after cloning a voice from a video with IndexTTS 2.5

1: choose IndexTTS 2.5, OmniVoice or AuK in the header. 2: the preset for your GPU is selected on the first start. 3: any audio or video file can be the voice. 4: type your script. 5: the result plays right away, with the elapsed time, speed and VRAM next to it.

1-Click Installation

What You Download

You get one small zip file with the installers for every platform.

The files inside Index_TTS_v8_1.zip

Extract it into a short folder path without spaces, double-click Windows_Install_or_Update.bat, then start the app with Windows_Start_App.bat. Run the same installer again at any time to update.

Latest PyTorch and CUDA 13

The installer makes its own Python 3.12 virtual environment, so your other apps stay untouched. It installs PyTorch 2.14.1 with CUDA 13, Transformers, Triton for Windows and every other package with uv, a very fast package installer.

Installer log: PyTorch 2.14.1 with CUDA 13 and 164 packages resolved in 7 seconds with uv

PyTorch, torchvision and torchaudio always come from the same CUDA 13 build. Our fresh install took 5 minutes 39 seconds on our PC, models included.

Automatic Model Downloads

At the end, the installer downloads the IndexTTS 2.5 models and their helper models by itself.

Model downloader: 16 connections per large file, SHA256 checks of the model weights, 22 model files and 3 helper models

Large files download over 16 connections, and every model weight file is checked with SHA256, so a broken download is caught right away. Running the installer again skips the large files you already have.

Start the App

Double-click Windows_Start_App.bat. On the first start, the app reads your GPU memory and selects the matching GPU preset.

First start: the 32 GB GPU preset is selected, the app is ready in 9.5 seconds, OmniVoice and the built-in Whisper download on first use

The app was ready in 9.5 seconds on the first start and in 4 to 6 seconds later. OmniVoice, AuK and the built-in Whisper download automatically the first time you use them.

Cloud GPUs: RunPod, SimplePod and Massed Compute

You can also run the app on a cloud GPU. The cloud installers set up everything with one command and download the models for you. A Gradio share link lets you use the app from any device.

The step-by-step commands are in the instruction files inside the zip.

Requirements

On Windows you need Python 3.12, Git, FFmpeg, CUDA 13, cuDNN 9.17 and Visual Studio with the C++ tools. This tutorial shows every step, and the same setup runs all our AI apps.

Three Speech Models in One App

Every model has its own strengths, and you switch between them with one click in the header.

Speech model list in the header: IndexTTS 2.5, OmniVoice and AuK

IndexTTS 2.5 gives you emotion control and reads five languages. OmniVoice clones voices, designs voices from tags and speaks many languages. AuK clones voices, designs voices from a plain sentence and edits recordings in its own tab. Each model keeps its own settings, and every preset stores all three.

GPU Presets for 6 GB to 32 GB Cards

Seven read-only presets fit the whole app to your card: generation, dataset preparation and training together. On the first start, the app selects the preset of your GPU.

Where to find it: the Universal preset list sits at the top of the page, above every tab.

Where to find the presets: Universal preset at the top of the page

Open the Universal preset list to see every preset:

Universal preset list: 7 GPU presets and the voice presets saved by trainings

Change anything you like, type a name and click Save to keep your own preset. It stores every setting of every tab and of all three models. Every training adds a ready-to-use preset of its new voice to this list. The Help tab shows what each GPU preset uses:

GPU VRAM presets table: peak memory, streamed GPT blocks, beams and diffusion steps per card size

A smaller card pays with speed first: helper models move to the CPU, and the 6 GB preset streams 22 of the 24 GPT blocks from RAM.

Voice Cloning From Any Audio or Video

A clean clip of a few seconds is enough to clone a voice. Audio and video files both work, and the preview plays exactly what the app will use.

Where to find it: open the Voice Generation tab. Reference Voice is the left column.

Where to find Reference Voice: the left column of the Voice Generation tab

Here is the Reference Voice column up close:

Reference Voice: upload, microphone recording, time ranges and local paths

Drop a file, record from your microphone, or load any file on your PC by its path. Time ranges such as 0.2:14.8 keep only the clean part of a long recording, and several ranges are joined in order.

Text, Smart Sentences and Pauses

Long text is split into sections before the voice reads it. Smart sentences packs whole sentences into each section.

Where to find it: in the Voice Generation tab, Text & Timing is the middle column.

Where to find Text and Timing: the middle column of the Voice Generation tab

Here are the text controls:

Text and Timing: an exact pause tag, whole sentences, and line length and pauses from a trained voice

Tags such as [pause:800ms] insert an exact silence. With a trained voice, the line length and the pauses between sentences come from the speaker's own recordings. Below the text, the live section preview shows every section before you generate:

Live section preview: 5 sections, an exact 800 ms pause, and the tokens and words of every section

You see the tokens and words of each section and every pause, so you can fix the text before you spend any GPU time.

Pronunciation Check and Dictionary

Names and technical words are the hardest part of any narration. The pronunciation check finds the words the model may not know and suggests a reading for each.

Where to find it: in the Voice Generation tab, click Pronunciation check & dictionary under the live section preview.

Where to find the pronunciation check: under the live section preview

Click Check unknown words:

Pronunciation check: unknown words with suggested readings and the pronunciation dictionary

Readings come from the CMU dictionary, acronym and CamelCase splitting, and letter rules. Save the good ones to your dictionary, and the app applies them in Voice Generation, Batch Generation and the live preview, for IndexTTS 2.5 and OmniVoice.

Take Quality: The Best Take of Every Section

Take quality picks the best take of every section. It renders several takes, lets the built-in Whisper check every word, and keeps the take that sounds most like the voice without a word error. It works for all three models, and the GPU presets switch it on.

Where to find it: in the Voice Generation tab, Take quality is in the right column, under the live log.

Where to find Take quality: the right column of the Voice Generation tab, under the live log

Here is Take quality after an OmniVoice clone of the same video clip:

Take quality: 10 takes of the section, the most similar take kept with no word errors

OmniVoice rendered 10 takes of our 32 second script in 26.5 seconds, and the app kept the most similar take with 0 word errors. In the developer's test on 60 lines of voice cloning, word errors fell from 0.65 % to 0.20 % with OmniVoice and from 0.91 % to 0.35 % with IndexTTS 2.5.

OmniVoice: Clone, Design or Auto Voice

OmniVoice clones a reference voice, designs a new voice from tags, or picks a voice by itself.

Where to find it: choose OmniVoice in the header. Its section opens in the Voice Generation tab, under Voice LoRA / DoRA.

Where to find the OmniVoice section: choose OmniVoice in the header, then its section in the Voice Generation tab

Here is the OmniVoice section with a designed voice:

OmniVoice section: voice modes, voice tags, supported tags and the sampling presets

We picked four tags: female, young adult, moderate pitch and british accent. The new voice spoke 6.1 seconds of speech in 1.2 seconds, with no reference recording at all.

AuK: Describe a Voice in Words

AuK clones voices, designs a voice from one plain sentence, and lets a trained AuK voice speak without a reference.

Where to find it: choose AuK in the header. Its section opens in the Voice Generation tab, under Voice LoRA / DoRA.

Where to find the AuK section: choose AuK in the header, then its section in the Voice Generation tab

Here is the AuK section with a described voice:

AuK section: voice modes, a voice description, ready examples and sampling presets

"A calm middle-aged man with a deep, warm voice, speaking slowly and clearly, like a documentary narrator." AuK made 6.6 seconds of speech from that sentence in 1.5 seconds. AuK speaks English and Chinese.

Emotion Control

IndexTTS 2.5 separates the voice from the emotion, and the app gives you four ways to set the emotion.

Where to find it: with IndexTTS 2.5 selected, scroll down in the Voice Generation tab and click Emotion Control. The Open / close all sections button at the top opens it too.

Where to find Emotion Control in the Voice Generation tab

Here is Emotion Control with the emotion sliders selected:

Emotion Control: 4 emotion sources, emotion weight, emotion description and 8 emotion sliders

Keep the speaker's own tone, copy the delivery of an emotion clip, mix joy, anger, sadness, fear, disgust, depression, surprise and calm, or describe the emotion in words. Emotion weight sets how strong it is.

Timing, Output Formats and Word Timestamps

Make the speech fit your video, save it in the format you need, and get subtitles of every word.

Where to find it: in the Voice Generation tab, click Segmentation & Timing, Output and Execution.

Where to find Segmentation and Timing, Output and Execution in the Voice Generation tab

The three sections opened:

Timing and output settings: target duration, speaking rate, WAV and MP3, word timestamps and isolated subprocess

Ask for an exact length, such as 25 seconds, and Natural mode paces the speech to fit. Every run saves a WAV file, one tick adds an MP3 up to 320k, and another adds SRT and VTT subtitles with the timing of every word. Isolated subprocess mode frees all VRAM after each job.

Subtitles to Timed Speech and MP4

Turn any subtitle file into speech that follows its timing, ready for dubbing, translations and slide videos.

Where to find it: in the Voice Generation tab, the Captions and Still image for MP4 boxes are under the text controls.

Where to find captions and the still image for MP4 in the Voice Generation tab

Drop your subtitle file and tick Use caption cue timing:

Captions: an SRT file with cue timing, a still image, and every cue in its time slot

SRT, VTT, SBV, ASS/SSA, SUB, LRC, TTML/DFXP, SAMI, JSON and TSV all work, and the app recognizes the format from the file itself. Add an image, and the same run also makes an MP4:

Finished MP4: the slide with the narration, 30.00 seconds with every cue on time

Our 6-cue subtitle file became exactly 30.00 seconds of speech, with every cue starting on time.

AuK Audio Editing

AuK edits real recordings: change the words, the delivery or the sound, and separate voices.

Where to find it: choose AuK in the header, then click the AuK Audio Editing tab.

Where to find AuK Audio Editing: choose AuK in the header, then the AuK Audio Editing tab

Here is a word replacement:

AuK Audio Editing: a recording, the task list, original and new words and the edited audio

We replaced "quick test" with "short demo" in a 9.9 second recording. The edit took 1.7 seconds, and Whisper heard the new words in the result. The 23 tasks also insert or remove words, change speed, pitch, volume and emotion, turn speech into a whisper, add a laugh, clean up noise and reverberation, and keep one speaker or the voices from music.

Batch Generation

Generate many scripts or subtitle files in one run.

Where to find it: click the Batch Generation tab.

Where to find Batch Generation: the Batch Generation tab

Drop your files, paste your texts or point the app at a folder, then click Generate batch:

Batch Generation: inputs, naming, reference mode and the results of three stories

Our three stories became 116 seconds of speech in 65 seconds with OmniVoice, every section picked from 10 takes. Use one voice for all files, or a matching audio file next to each script for a different voice per file. Continue after item errors keeps the batch going.

Train Your Own Voice

A trained voice follows the speaker much more closely than a single reference clip: the pace, the pauses and the way each word is said. You can train IndexTTS 2.5, OmniVoice and AuK, and the app takes you from videos and subtitles to a finished voice.

Prepare the Dataset

Point the app at your videos or audio files with their subtitles. It cuts them into clean clips of whole sentences.

Where to find it: click the LoRA Dataset Preparation tab.

Where to find dataset preparation: the LoRA Dataset Preparation tab

Here is a real run on four of our tutorial recordings:

Dataset preparation: inputs, the default settings, 4 recordings and the finished run with 205 clips

Your subtitles give the words and Whisper gives the exact word timing. The four recordings, 46.8 minutes in total, became 205 clean clips of 40.6 minutes in 50 seconds. Every clip can be checked before training:

Prepared clips: statistics, the clip length chart and every clip with its text

The clips are 4 to 16 seconds long, and every clip keeps its text.

Voice and Transcript Audit

The audit checks every clip against a clean recording of the speaker and against its text, and builds a separate training dataset from the clips that pass.

Where to find it: in the LoRA Dataset Preparation tab, click Voice and transcript audit at the bottom.

Where to find the voice and transcript audit: the bottom of the LoRA Dataset Preparation tab

Here is the audit of our 205 clips:

Voice and transcript audit: speaker reference, held-out recordings, limits and the result: 194 clips kept

It checked 205 clips in 1 minute 14 seconds and kept 194. A second Whisper model rescued 7 good clips that the first one misheard. A whole recording can be held out, so the app tests the voice on speech it never trained on.

Start the Training

Choose your dataset, give the voice a name and click Start training. The defaults are the settings we measured on real voices.

Where to find it: click the Voice Training tab.

Where to find training: the Voice Training tab

Here is the tab with our audited dataset:

Training setup: audited dataset, GPU VRAM preset, fluency filter, training method and the automatic checks

The GPU VRAM preset fits the training to your card. DoRA and LoRA train small adapters; full fine-tuning trains the speech model's own weights on 16 GB cards and larger (AuK 32 GB). The fluency filter can train on fluent clips only. During training the dashboard updates live:

Live training dashboard: progress, speed and VRAM, epoch checks and the loss, learning rate, gradient and speed charts

Our IndexTTS 2.5 voice trained at about 5 to 6 steps per second with 2.3 GB of VRAM. The app checks the voice on held-out clips every epoch and keeps the best checkpoint.

Automatic Checks After Training

When training ends, the app keeps working for you. It renders the same sentences with the base model and with the checkpoints, then compares them with the speaker's real recordings of those sentences. It also calibrates the speaking rate, trains a voice decoder adapter, tunes the decoding settings and auditions reference clips. Here is the comparison for our voice:

Automatic speech comparison: every checkpoint against the base model on real speech, and the chosen voice with its decoder adapter

Likeness to the speaker's real recordings rose from 0.810 with the base model to 0.867 for the chosen checkpoint with its voice decoder adapter, and the word error rate fell from 2.4 % to 2.3 %. The app picked that checkpoint by itself and set it as the voice's default.

A Ready-to-Use Preset

Every training ends with a ready-to-use preset of the new voice: our training saved my_voice_Takes_5 in the Universal preset list you saw above. It loads the chosen checkpoint, the best reference clip from the audition, the calibrated speaking rate of 0.86, the decoding settings from the sweep, and Take quality with up to 5 takes per section.

Use Your Trained Voice

Select your voice in Voice Generation, and the app loads everything it learned.

Where to find it: in the Voice Generation tab, the Voice LoRA / DoRA panel is right under the three columns. You can also load the voice's ready-made preset from the Universal preset list.

Where to find the trained voice panel: Voice LoRA / DoRA under the three columns

Here is the panel after loading the preset of our new voice:

Voice LoRA / DoRA panel: trained voice, decoder adapter, calibrated pace, line length, the speaker's pauses and decoding

The panel shows the voice's calibrated pace, the line length and the pauses measured from its own training clips, and the decoding its sweep adopted, and applies them to your text. With this preset our voice read the 88-word script as 35.2 seconds of speech in 28 seconds, and the built-in Whisper found 0 word errors in every section.

Compare Checkpoints by Ear

The Checkpoint Grid tab renders the same texts with every checkpoint you choose, so you can listen to them side by side.

Where to find it: click the Checkpoint Grid tab.

Where to find checkpoint comparison: the Checkpoint Grid tab

Here is a grid of our voice with the base model and three checkpoints:

Listening grid: the base model and three checkpoints reading the same texts with the same reference and seed

Every cell uses the same texts, reference and seed, so you hear only the difference between the checkpoints. Every grid is saved and opens again at any time, and Calibrate speaking rate from this grid measures the voice's real pace.

Models and Performance

Runtime

Fit the models to your GPU in one place.

Where to find it: click the Models & Performance tab.

Where to find the runtime settings: the Models and Performance tab

Here are the runtime controls:

Runtime: VRAM tier, fit check, apply or free VRAM, BF16 or INT8 ConvRot and the attention backend

Choose a VRAM tier, and the fit check shows the expected peak memory before anything loads. Pick the BF16 model or the INT8 ConvRot model, which saves memory. Unload model / free VRAM clears the GPU in one click.

Block Swap and Model Placement

Small cards can stream parts of the model from system RAM.

Where to find it: in the Models & Performance tab, click Block Swap & Memory and Auxiliary Model Residency.

Where to find block swap and model placement in the Models and Performance tab

Here is the 6 GB tier:

6 GB tier: 22 of 24 GPT blocks streamed from RAM and the placement of every helper model

The 6 GB tier streams 22 of the 24 GPT blocks from RAM and keeps the reference models on the CPU. Every helper model can stay on the GPU, stay on the CPU or load on demand.

Model Files and VRAM Benchmark

Check your model files and test a GPU tier before a long job.

Where to find it: in the Models & Performance tab, click Model Files & Downloads and VRAM Benchmark.

Where to find model files and the VRAM benchmark in the Models and Performance tab

Both sections after a real test:

Model files with the INT8 ConvRot GPT, and a VRAM benchmark of the 8 GB tier with its memory cap

The INT8 ConvRot GPT is 1.1 GB instead of 3.1 GB and downloads from inside the app. The VRAM benchmark ran the 8 GB tier with its memory limit in place on our card: it fit with a 3.5 GB peak and made 25.9 seconds of speech in 19.3 seconds.

Built-in Help and Changelog

You never need to leave the app to learn it.

Where to find it: click the Help tab or the Changelog tab.

Where to find help: the Help and Changelog tabs

The Help tab starts with a four-step quick start:

Help tab: four-step quick start and every speech model explained

It explains every model, every workflow and every parameter. The Changelog tab lists every release in plain words:

Changelog tab: every release with its changes in plain words

The newest release is always at the top.

Latest Updates

The app gets frequent updates. Versions 7.0 to 8.1 came out between 2 and 4 October 2026:

  • 8.1: fresh installs keep PyTorch, torchvision and torchaudio on one CUDA 13 build, the preset after training carries the voice's calibrated speaking rate and decoding, IndexTTS training caches its features by itself, and the chosen speech model stays selected after a reload.
  • 8.0: Take quality for all three models, a ready-to-use preset after every training, and OmniVoice as the default model of a new install.
  • 7.1: AuK, the third speech model, with voice design from a sentence, trained voices without a reference, and the AuK Audio Editing tab.
  • 7.0: OmniVoice, the second speech model, with voice design, BF16 or INT8 ConvRot and training.

Get Ultimate Text To Speech Generator

Download the latest zip file attached to this post, extract it and run the installer for your platform. To update, get the newest zip, overwrite the old files and run Windows_Install_or_Update.bat again.

Top comments (0)