This guide turns the video tutorial into a step-by-step reference for installing Voicebox, creating a voice profile, generating local AI speech, building multi-voice stories, applying effects, and getting cleaner results — with an important consent-first approach to voice cloning.
The video describes Voicebox as a local-first AI voice studio: voice data is processed on your own computer rather than uploaded to a cloud service. That makes it useful for privacy-sensitive voice work and experimentation.
Content creators, developers, students, educators, storytellers, game/dialogue creators, and anyone who wants to experiment with local AI speech generation.
| Feature | What it does | Best use |
|---|---|---|
| Voice profiles | Pair a voice sample with a transcript and settings. | Reusable narration, characters, or approved personal voices. |
| TTS engines | Generate speech using different models with different strengths. | Trying quality, speed, multilingual support, or expressiveness. |
| Stories editor | Multi-track timeline for dialogue and podcast-like scenes. | Conversations, audiobook dialogue, game characters, demos. |
| Effects | Post-process generated voices with presets or custom chains. | Sci-fi voices, radio comms, echo spaces, lower/deeper voices. |
| GPU acceleration | Uses available GPU/Metal/CUDA where supported. | Faster generation on compatible hardware. |
Voice cloning can be useful, creative, and accessible — but it can also be misused. The video explicitly notes that the process should be used for your own voice, a character you created, or someone who has given explicit consent. Treat this as a hard rule.
For public or semi-public work, consider adding: This audio includes AI-generated voice content created with permission.
The video directs viewers to voicebox.sh. From there, find the download section and choose the correct installer for your operating system.
| Platform | What to choose | Notes |
|---|---|---|
| Windows | Windows 64-bit MSI | Run the installer and follow the prompts. |
| Mac Apple Silicon | Apple Silicon / ARM download | Use this for M1, M2, M3, or M4 Macs. |
| Mac Intel | Intel Mac option | Check Apple menu → About This Mac if unsure. |
| Linux | Build-from-source guide | The video references voicebox.sh/linux-install. |
The first launch may set up files or models in the background. Let it finish before judging performance.
A voice profile is the reusable identity Voicebox uses for generation. It typically includes a clean audio sample, a transcript, a name, and optional personality/default-engine settings.
Use the left sidebar microphone/voices icon. Click New Voice or Create Voice.
| Input option | Use when | Tip |
|---|---|---|
| Upload audio | You already have an MP3/WAV or other recording. | Use clean, dry speech without music. |
| Microphone | You want to record directly in the app. | Aim for 20–30 seconds in a quiet room. |
| System audio | You need to capture audio playing on your computer. | Use only audio you have rights/permission to capture. |
Find a quiet spot. Avoid fans, music, keyboard noise, and other people speaking. Talk naturally — not too fast, not over-enunciated. Say your name, describe what you are doing, or tell a short story.
Click Transcribe. The video says Voicebox may download a Whisper model the first time. After transcription, check the text for major mistakes because the text/audio pair helps create the voice profile.
Give the voice a clear name. Optionally describe its personality or speaking style. Choose a default engine or leave it as no preference if you want to experiment later. Click Create Profile.
Select the voice profile you created. Enter the text you want the voice to say in the text box.
The video mentions several engines and describes different strengths. Engine names may evolve as the app updates, so treat this as a starting map rather than a permanent list.
| Engine mentioned | Video’s description | Try it for |
|---|---|---|
| Qwen/“Quan” 3 TTS | High-quality multilingual speech; can accept instructions like “speak slowly” or “sound excited.” | Narration, multilingual work, quality tests. |
| Chatterbox | Good language coverage; the video says it covers 23 languages. | Languages beyond English. |
| Chatterbox Turbo | Fast and expressive; supports emotion tags in text. | Dialogue, quick iteration, expressive lines. |
| Kokoro | Tiny and fast; includes many preset voices. | Older hardware or very fast drafts. |
For your first generation, use one or two sentences. Example:
Leave effects off for the first run so you can evaluate the raw voice quality.
Click Generate. The first run may take longer if a model needs to download. Listen to the output. If it is close but not perfect, regenerate two or three times; different takes can vary noticeably.
Use the three-dot menu next to a generation and choose export. The video says Voicebox can export WAV or MP3 for videos, presentations, podcasts, stories, or other projects.
The Stories editor is the feature that turns Voicebox from a single-voice generator into a small production studio. Each generated line becomes a clip on a timeline.
Use the left sidebar icon that looks like a timeline or stacked layers. Click New Story, name it, and create it.
Type the first line in the text field, choose the speaker/voice profile, and click generate. The line appears as a clip.
Alternate between voice profiles for a podcast-style conversation, fictional dialogue, training scene, or game script.
| Speaker | Example line | Production note |
|---|---|---|
| Host | “Hey, did you hear about that new AI tool everyone’s talking about?” | Use a clear, upbeat voice. |
| Guest | “Yeah, it’s called Voicebox. I’ve been using it all week.” | Use a second profile or preset voice. |
| Narrator | “The two creators opened the app and started building a story.” | Add neutral pacing between dialogue sections. |
Drag clips to reorder them, adjust timing, preview the full conversation, then export the completed story as a single audio file.
You can also build custom chains with effects such as reverb, delay, chorus, compressor, and pitch shift.
The video says Voicebox can use hardware acceleration where supported:
Get the voice sounding good first, then add effects. Effects can enhance a good generation, but they rarely fix a noisy, poorly cloned, or badly paced one.
Background noise is the fastest way to reduce quality. Close windows, move away from fans, avoid music, and record in a quiet room.
The video recommends getting close to 30 seconds if possible. More clean data generally gives the model a better sense of the voice.
Do not rush, mumble, or over-act the sample. Natural pacing helps the clone sound less artificial.
One engine may be best for narration while another is better for expressive dialogue or multilingual speech.
If a take is close but not quite right, regenerate a few times before changing everything. The video notes that takes can vary because of random seeds.
Short test lines make it faster to compare engines, voices, and effects before producing a long script.
| Problem | Likely cause | Fix |
|---|---|---|
| Voice sounds noisy or unstable | Bad source recording, background noise, or too little sample audio. | Record again in a quieter room; aim for 20–30 seconds of clean speech. |
| First generation is slow | Model download or no acceleration enabled. | Wait for the first setup to finish; check GPU settings afterward. |
| Clone does not sound like the original | Engine mismatch, poor transcript, or sample quality issue. | Correct the transcript, try Chatterbox Turbo/Qwen/Kokoro alternatives, regenerate several times. |
| macOS blocks the app | Unverified developer warning. | System Settings → Privacy & Security → Open Anyway, if you trust the source. |
| Windows warns about the installer | New open-source app not recognized by Defender. | Verify you downloaded from the official site/repo, then use More info → Run anyway if comfortable. |
| Story timing feels awkward | Clips need spacing or reordering. | Drag clips in the timeline, adjust pauses, and preview before exporting. |
Voicebox is presented in the video as a free, local-first AI voice studio that can handle voice profiles, speech generation, story timelines, effects, and hardware acceleration. The best beginner path is simple: install it, record a clean consent-based sample, create a profile, generate a short test, compare engines, then move into stories and effects once the basic voice sounds good.
Best first project: create a 20-second personal voice sample, generate a two-line narration, export it, then repeat with a second engine to compare quality.