DEV Community

orca forge
orca forge

Posted on Originally published at forge.workstyle.tech

We've Released Voice Canvas as OSS: Easily Create Your Favorite Voice

📝 Originally published (in Japanese) at forge.workstyle.tech.

I created "Voice Canvas," a tool that lets you define voice characteristics using sliders to convert original audio into any voice you choose.
My goal is to make it easy for anyone to create the voices they like.
The code is released as an open-source project under the MIT license. It is currently in beta.
It runs without a GPU, though conversion on a CPU takes between one to several minutes per execution.

Process of creating and converting voices with Voice Canvas

Watch a digest video of the operations (3 minutes, with audio)

What You Can Do with Voice Canvas

Voice Canvas is an app that changes voice timbre while preserving the original audio's speech content, tempo, and intonation. You configure eight vocal characteristics—such as pitch and warmth—using sliders, or adjust them directly by dragging points on a graph.

There are three ways to provide source audio: generate speech from text using VOICEVOX, record live from a microphone, or import an audio file. If you don't have an audio file ready, you can use the bundled sample clips. (VOICEVOX is a text-to-speech engine; since it isn't bundled with Voice Canvas, you'll need to install it separately if you plan to use TTS.)

UI showing source audio, eight sliders, and the new voice side by side

When you adjust the voice settings and convert, each result is listed on the right as a "take." A take is an individual conversion result. Because you can play them back while viewing the underlying settings, it's easy to compare different voices generated from the exact same source audio. In the finishing stage, you can fine-tune speech speed and pitch fluctuations before exporting to WAV. We made sure these finishing touches can be reapplied without having to re-run the full conversion.

Out of the box, it comes with 69 "anchors" as voice building blocks. Anchors are reference audio recordings of real speakers that serve as models for crafting voices. Voice Canvas automatically selects up to four speakers closest to your slider settings and blends their voices together. You can also add your own recordings or audio files, as long as they are at least 8 seconds long.

Additionally, the built-in media tools allow you to extract audio from video, remove background music and noise, and trim or merge audio clips. You can also use these utility features to prep and clean up audio before registering it as a new anchor.

How Voice Ingredients are Blended

When you move the "Age" or "Warmth" sliders, Voice Canvas searches for anchors close to those values. For age, it uses the speaker's age group label, and for gender, it doubles the weight when calculating proximity, ultimately selecting up to four top-ranked anchors.

The mechanism for creating voices by blending multiple anchors

Next, Seed-VC transforms the voice quality using the blended vector of the selected speakers along with a 12-second audio clip from each as a reference. Finally, you can adjust the speaking speed and pitch fluctuations. These adjustments can be made without re-running the conversion. The output audio is in WAV format.

Since we use the voices of real people, we have taken great care in how the materials are handled. The package includes 45 speakers from Common Voice Japanese (which has clear terms of use) and 24 speakers from JVNV (Japanese Natural Emotional Speech Corpus), consisting of 4 speakers across 6 emotions. The former is CC0, while the latter is CC BY-SA 4.0. JVNV requires attribution.

Challenges and Design Decisions

How to Turn Slider Values into a Voice

The first hurdle was figuring out how to connect slider values to actual voice conversion. You can't just directly map raw numbers to an audio output. Instead, I adopted an approach that uses "anchors"—reference samples with distinct vocal characteristics—as building blocks.

The system automatically selects anchors closest to the slider settings, blends up to four of them together, and passes the result to Seed-VC. Users don't need to manually pick each reference voice one by one; they can simply adjust the sliders and hit convert to explore different vocal profiles. You can also add your own anchors, which naturally expands the variety of voices you can generate based on your reference materials.

I haven't quantified exactly "how many unique voices" this system can create. The sliders take continuous values, but I haven't formally benchmarked similarity scores or clustering to provide a hard number. For now, the intended workflow is hands-on: tweak the settings, generate takes side by side, and compare them by ear.

It Runs on CPUs—at the Cost of Conversion Speed

I made sure the tool works without a dedicated GPU. That said, CPU-based conversion is heavy. Benchmarking in my development environment with the default model and default settings showed that a 3.5-second clip took about 2 minutes, and a 10-second clip took roughly 3.5 minutes. Increasing thread counts didn't improve speeds either.

Reducing the step count brought a 2.5-second clip down to about 70 seconds, though I haven't evaluated the resulting drop in audio quality. For the walkthrough in this article, we'll stick to the defaults first. Once you trigger the first conversion, an elapsed timer appears on the right side, so just sit tight. Being CPU-compatible doesn't mean it's fast—setting realistic expectations about wait times up front felt essential for first-time users.

Rethinking the Criteria for Registering Anchors

When adding new anchors, the tool inspects audio length and recording quality. Initially, I enforced strict validation rules—such as requiring at least 15 seconds of audio and high fidelity—and rejected anything that fell short. However, this made the barrier to entry way too high.

To fix this, I loosened the outright rejection criteria to only trigger if the audio is under 8 seconds or contains no detectable speech. Even if an audio sample fails to hit optimal quality guidelines or voice isolation checks, the tool still registers it and flags a warning. Users can check the reason behind the warning and clean up background music or noise using external audio tools if needed. By narrowing hard failures, users have the agency to decide for themselves.

Connecting from Windows/WSL2 to VOICEVOX

To ensure Windows compatibility, I tested everything on Windows 11 using WSL2 (Windows Subsystem for Linux). Initially, Voice Canvas running inside WSL couldn't connect to VOICEVOX running on the Windows host. Under WSL2's default networking mode, 127.0.0.1 inside WSL doesn't route to localhost on the Windows side.

I added instructions in the README on how to enable WSL's mirrored networking mode. You can configure this by adding settings to .wslconfig in your Windows user profile folder and restarting WSL. I'll provide an example configuration later in this guide. Of course, you don't have to use VOICEVOX—you can always get started with the built-in sample clips, microphone recordings, or local audio files.

Streamlining the First-Run Experience

If you launch the app without any source audio or reference material, you're immediately stuck. To prevent this, the app automatically extracts 69 bundled anchor voices on its first run and includes sample clips out of the box. If VOICEVOX isn't detected, it prompts you with alternatives like mic recording, uploading files, or using sample audio.

I evaluated three layout concepts before settling on a three-column workbench layout: source audio on the left, voice parameters in the middle, and the converted voice on the right. This left-to-right flow makes it immediately obvious what step to take next.

A significant portion of development was offloaded to Claude Code and Codex. Codex drafted the initial versions of the README and documentation, while UI copy and manuals were polished using yomiyasu. My role was defining specifications and asset validation criteria, testing implementations, and tying everything together.

Installing and Creating Your First Voice

From here, we will guide you through launching the app and converting a sample voice. First, let's start with a sample voice without using a text-to-speech software, and we can prepare VOICEVOX later.

Things to Know in Advance

The confirmed operating environments are Mac, Linux, and Windows 11 with WSL2 (Ubuntu 22.04). It does not work with Windows PowerShell as described in the README. For Windows, please open WSL2's Ubuntu and execute the commands inside it.

You will need Python 3.10, Node.js 20 or later, and git. Python's venv is a mechanism to separate dependent packages for this app. When you start it for the first time, it will automatically obtain a model of about 3.2GB. If you use BGM removal, you will need an additional approximately 80MB. Please keep your network connection until the download is complete.

Launching on Mac and Linux

Open a terminal and get the repository with the following command:

git clone https://github.com/maccotaro/voice-canvas.git voice-canvas && cd voice-canvas
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the terminal's working directory is voice-canvas, you are ready.

Next, get Seed-VC (the voice quality conversion mechanism).

git clone https://github.com/Plachtaa/seed-vc external/seed-vc
Enter fullscreen mode Exit fullscreen mode

Confirmation: If external/seed-vc is created and the acquisition process is completed, it's OK.

Create a virtual environment with Python 3.10.

python3.10 -m venv .venv-seedvc
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the command ends without errors, it's OK. venv refers to the Python environment for this project.

Install the necessary packages.

.venv-seedvc/bin/pip install -r service/voice-canva/requirements.txt
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the installation proceeds to the end and the command ends, it's OK. This may take some time.

Start the inference service. The inference service is the part that performs the voice conversion calculation.

bash scripts/start-backend.sh
Enter fullscreen mode Exit fullscreen mode

Confirmation: The first time, the model acquisition will start. After completion, open http://127.0.0.1:8770/health from another terminal to check.

Open another terminal, move to the same repository folder, and start the web screen.

bash scripts/start-web.sh
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the startup message is displayed, open http://localhost:3010 in your browser. If the Voice Canvas screen is displayed, the startup is complete.

Launching on Windows 11 with WSL2

WSL2 is a mechanism that runs a Linux environment such as Ubuntu on Windows. We have confirmed that it works on Windows 11 and WSL2 (Ubuntu 22.04). First, open Ubuntu. The following commands are entered into the Ubuntu terminal.

Install the necessary packages.

sudo apt update
sudo apt install -y git build-essential python3.10-venv python3.10-dev
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the installation is complete and you return to the prompt, it's OK. python3.10-venv is for the virtual environment, and build-essential and python3.10-dev are used for building some packages.

Also, prepare Node.js 20 or later inside WSL. Sometimes it won't start if the Windows-side Node.js gets mixed in, so make sure which one you are using.

which node
Enter fullscreen mode Exit fullscreen mode

Confirmation: If the displayed path is the one in the Linux of WSL, it's OK. If the Windows-side path comes out, set it to use the Node.js inside WSL.

The subsequent repository acquisition, Seed-VC acquisition, virtual environment creation, package installation, and backend and web startup are the same as the Mac and Linux commands. All are executed in the Ubuntu terminal.

If you use the Windows version of VOICEVOX to read articles, you will also need to set up the network on the WSL side. Create a .wslconfig file in the Windows user folder and write the following:

[wsl2]
networkingMode=mirrored
Enter fullscreen mode Exit fullscreen mode

Stop WSL with PowerShell and then reopen Ubuntu.

wsl --shutdown
Enter fullscreen mode Exit fullscreen mode

Confirmation: If you can connect to VOICEVOX from Voice Canvas started in WSL, it's OK. wsl --shutdown will stop the work in progress in WSL, so save it first. If you can't connect, press the "connect" button according to the app's "VOICEVOX not found" instructions. If you don't use VOICEVOX, you can proceed with the sample voice.

Creating Your First Take

Open http://localhost:3010 in your browser. At first, the included anchors will be automatically expanded. If you have 69 usable anchors for voice settings, you are ready.

First, select a sample voice from "Original Voice". If you have a recorded audio, you can also import the audio file. When reading an article, start VOICEVOX, enter the article, and press "Read aloud".

Confirmation: If the left audio column says "OK", you have prepared the original voice. If it says "Not set", try selecting the sample voice again.

In the central voice settings, move the sliders of the 8 characteristics that you care about. At first, try changing only one to make it easier to compare the settings and results.

Confirmation: If the sliders and radar chart move in sync, it's OK. You can adjust the voice settings as many times as you want before conversion.

Press "Convert to this voice".

Confirmation: If the processing time is displayed on the right side in "New Voice" and the processing is complete, and one take is added, it's a success. It may take several minutes on the CPU. Wait without closing the screen.

Select a take and play it back. If you convert again with different slider values, another result will be added.

Confirmation: If multiple takes are lined up on the right side and you can play each one back, the comparison is complete. You can choose the take you like and save it as a WAV file from "Finish and export".

Troubleshooting

If the convert button is disabled, please check the original audio field on the left. If it says "Unset" (未設定), no audio has been loaded yet. When you change the text or the speaker, the previous voice will no longer be available, so please click "Read aloud" again.

If the message "Insufficient anchors" appears, you have fewer than three anchors currently in use. Either return anchors you removed in Anchor Management or add a new voice.

If VOICEVOX cannot be found, click "Connect" while keeping VOICEVOX open. If you are running Voice Canvas on Windows WSL2, check the .wslconfig settings mentioned above. If you wish to continue without using text-to-speech, you can select recording, audio file, or sample audio.

If you cannot use the microphone, please allow the browser to access your microphone. If a video fails to load, verify that it is in a format supported by your browser. Audio in the asset tray persists when moving between pages, but it will disappear if you refresh the page. Please save any audio you wish to keep to a card.

Detailed instructions are also available in the manual, which you can open via "How to Use" (使い方) at the top right of the app.

Licensing and Assets

The Voice Canvas code is licensed under the MIT License. Seed-VC is licensed under GPL-3.0 and is not included in the repository; it is fetched at runtime. If you distribute a Docker image containing Seed-VC, please comply with the GPL-3.0 license. Model weights are also not included in the repository and are automatically fetched on first use.

The included anchors are sourced from Common Voice and JVNV. When modifying or redistributing JVNV materials or the sample audio created using them, follow the CC BY-SA 4.0 license and provide proper attribution. For Common Voice, ensure you do not identify individual speakers. If using VOICEVOX as the original voice, credit the character used and review each character's terms of use.

When adding voices, register audio that you have permission to use, such as recordings with the speaker's consent. Analyzed data that can be resynthesized is not stored; only 12-second reference audio, speaker vectors, and statistical values are handled.

Voice Canvas is in beta. For detailed usage and licensing information, refer to the repository's README and manual.

Repository: https://github.com/maccotaro/voice-canvas

Top comments (0)