Published 2026-07-15 • Based on AI BrainBox’s video “This Free AI Voice Generator is Insane”

Voicebox Local AI Voice Generator: Beginner’s Guide

This guide turns the video tutorial into a step-by-step reference for installing Voicebox, creating a voice profile, generating local AI speech, building multi-voice stories, applying effects, and getting cleaner results — with an important consent-first approach to voice cloning.

Watch the source video

Voicebox Workflow Install local app Record or upload clean voice sample Transcribe + create voice profile Generate, edit, export, and publish audio

What Voicebox is

Local-first AI voice studio

The video describes Voicebox as a local-first AI voice studio: voice data is processed on your own computer rather than uploaded to a cloud service. That makes it useful for privacy-sensitive voice work and experimentation.

Main capabilities

  • Clone a voice from a short sample.
  • Generate speech in multiple languages/engines.
  • Create reusable voice profiles.
  • Build multi-voice podcast or story scenes.
  • Apply effects such as robotic, radio, echo, and deep voice.
  • Export finished audio as WAV or MP3.

Who it is for

Content creators, developers, students, educators, storytellers, game/dialogue creators, and anyone who wants to experiment with local AI speech generation.

Feature map from the video

FeatureWhat it doesBest use
Voice profilesPair a voice sample with a transcript and settings.Reusable narration, characters, or approved personal voices.
TTS enginesGenerate speech using different models with different strengths.Trying quality, speed, multilingual support, or expressiveness.
Stories editorMulti-track timeline for dialogue and podcast-like scenes.Conversations, audiobook dialogue, game characters, demos.
EffectsPost-process generated voices with presets or custom chains.Sci-fi voices, radio comms, echo spaces, lower/deeper voices.
GPU accelerationUses available GPU/Metal/CUDA where supported.Faster generation on compatible hardware.

Consent and safety first

Only clone voices you have permission to use

Voice cloning can be useful, creative, and accessible — but it can also be misused. The video explicitly notes that the process should be used for your own voice, a character you created, or someone who has given explicit consent. Treat this as a hard rule.

Good uses

  • Your own narration voice for videos or presentations.
  • Fictional characters you created.
  • Authorized employee training voices.
  • Accessibility or assistive communication with permission.
  • Prototype podcast/story dialogue.

Avoid

  • Impersonating real people without consent.
  • Political, financial, legal, or medical deception.
  • Uploading or distributing someone else’s private recording.
  • Using cloned audio without disclosure where listeners could be misled.

Practical disclosure line

For public or semi-public work, consider adding: This audio includes AI-generated voice content created with permission.

1. Install Voicebox

1Go to the official site

The video directs viewers to voicebox.sh. From there, find the download section and choose the correct installer for your operating system.

2Choose your platform

PlatformWhat to chooseNotes
WindowsWindows 64-bit MSIRun the installer and follow the prompts.
Mac Apple SiliconApple Silicon / ARM downloadUse this for M1, M2, M3, or M4 Macs.
Mac IntelIntel Mac optionCheck Apple menu → About This Mac if unsure.
LinuxBuild-from-source guideThe video references voicebox.sh/linux-install.

3Run the installer

  • Windows: open the MSI. If Windows Defender warns that the app is unrecognized, the video says to use “More info” → “Run anyway” for this open-source app.
  • Mac: open the DMG, drag Voicebox into Applications, then open it. If macOS blocks it, go to System Settings → Privacy & Security → Open Anyway.

4Launch Voicebox

The first launch may set up files or models in the background. Let it finish before judging performance.

2. Create your first voice profile

A voice profile is the reusable identity Voicebox uses for generation. It typically includes a clean audio sample, a transcript, a name, and optional personality/default-engine settings.

Open Voices
New Voice
Record or upload
Transcribe
Create profile

1Open the Voices section

Use the left sidebar microphone/voices icon. Click New Voice or Create Voice.

2Choose an input method

Input optionUse whenTip
Upload audioYou already have an MP3/WAV or other recording.Use clean, dry speech without music.
MicrophoneYou want to record directly in the app.Aim for 20–30 seconds in a quiet room.
System audioYou need to capture audio playing on your computer.Use only audio you have rights/permission to capture.

3Record a clean sample

Find a quiet spot. Avoid fans, music, keyboard noise, and other people speaking. Talk naturally — not too fast, not over-enunciated. Say your name, describe what you are doing, or tell a short story.

4Transcribe the sample

Click Transcribe. The video says Voicebox may download a Whisper model the first time. After transcription, check the text for major mistakes because the text/audio pair helps create the voice profile.

5Name and save the profile

Give the voice a clear name. Optionally describe its personality or speaking style. Choose a default engine or leave it as no preference if you want to experiment later. Click Create Profile.

3. Generate speech from your profile

1Open the Generate tab

Select the voice profile you created. Enter the text you want the voice to say in the text box.

2Choose a TTS engine

The video mentions several engines and describes different strengths. Engine names may evolve as the app updates, so treat this as a starting map rather than a permanent list.

Engine mentionedVideo’s descriptionTry it for
Qwen/“Quan” 3 TTSHigh-quality multilingual speech; can accept instructions like “speak slowly” or “sound excited.”Narration, multilingual work, quality tests.
ChatterboxGood language coverage; the video says it covers 23 languages.Languages beyond English.
Chatterbox TurboFast and expressive; supports emotion tags in text.Dialogue, quick iteration, expressive lines.
KokoroTiny and fast; includes many preset voices.Older hardware or very fast drafts.

3Start with a short test

For your first generation, use one or two sentences. Example:

Leave effects off for the first run so you can evaluate the raw voice quality.

4Generate, listen, and regenerate

Click Generate. The first run may take longer if a model needs to download. Listen to the output. If it is close but not perfect, regenerate two or three times; different takes can vary noticeably.

5Export finished audio

Use the three-dot menu next to a generation and choose export. The video says Voicebox can export WAV or MP3 for videos, presentations, podcasts, stories, or other projects.

4. Build multi-voice stories and podcasts

The Stories editor is the feature that turns Voicebox from a single-voice generator into a small production studio. Each generated line becomes a clip on a timeline.

1Open Stories

Use the left sidebar icon that looks like a timeline or stacked layers. Click New Story, name it, and create it.

2Add the first line

Type the first line in the text field, choose the speaker/voice profile, and click generate. The line appears as a clip.

3Add more lines with different voices

Alternate between voice profiles for a podcast-style conversation, fictional dialogue, training scene, or game script.

SpeakerExample lineProduction note
Host“Hey, did you hear about that new AI tool everyone’s talking about?”Use a clear, upbeat voice.
Guest“Yeah, it’s called Voicebox. I’ve been using it all week.”Use a second profile or preset voice.
Narrator“The two creators opened the app and started building a story.”Add neutral pacing between dialogue sections.

4Arrange and export

Drag clips to reorder them, adjust timing, preview the full conversation, then export the completed story as a single audio file.

5. Use effects and hardware acceleration

Built-in effects mentioned

  • Robotic digital/sci-fi character sound.
  • Radio walkie-talkie or vintage broadcast feel.
  • Echo chamber reverb, cave, or large-space feel.
  • Deep voice lower pitch and heavier tone.

You can also build custom chains with effects such as reverb, delay, chorus, compressor, and pitch shift.

GPU acceleration

The video says Voicebox can use hardware acceleration where supported:

  • Windows + Nvidia: CUDA support can be set up from inside the app.
  • Mac M-series: Metal acceleration is used by default.
  • Dedicated GPUs from Nvidia, AMD, or Intel Arc may improve generation speed.

Effects rule of thumb

Get the voice sounding good first, then add effects. Effects can enhance a good generation, but they rarely fix a noisy, poorly cloned, or badly paced one.

6. Pro tips for better results

Record clean audio

Background noise is the fastest way to reduce quality. Close windows, move away from fans, avoid music, and record in a quiet room.

Aim for 20–30 seconds

The video recommends getting close to 30 seconds if possible. More clean data generally gives the model a better sense of the voice.

Speak naturally

Do not rush, mumble, or over-act the sample. Natural pacing helps the clone sound less artificial.

Try multiple engines

One engine may be best for narration while another is better for expressive dialogue or multilingual speech.

Regenerate

If a take is close but not quite right, regenerate a few times before changing everything. The video notes that takes can vary because of random seeds.

Keep first tests short

Short test lines make it faster to compare engines, voices, and effects before producing a long script.

7. Troubleshooting

ProblemLikely causeFix
Voice sounds noisy or unstableBad source recording, background noise, or too little sample audio.Record again in a quieter room; aim for 20–30 seconds of clean speech.
First generation is slowModel download or no acceleration enabled.Wait for the first setup to finish; check GPU settings afterward.
Clone does not sound like the originalEngine mismatch, poor transcript, or sample quality issue.Correct the transcript, try Chatterbox Turbo/Qwen/Kokoro alternatives, regenerate several times.
macOS blocks the appUnverified developer warning.System Settings → Privacy & Security → Open Anyway, if you trust the source.
Windows warns about the installerNew open-source app not recognized by Defender.Verify you downloaded from the official site/repo, then use More info → Run anyway if comfortable.
Story timing feels awkwardClips need spacing or reordering.Drag clips in the timeline, adjust pauses, and preview before exporting.

Quick reference

First voice checklist

  • Download Voicebox from official source.
  • Install and launch the app.
  • Open Voices → New Voice.
  • Record/upload a clean 20–30 second sample.
  • Transcribe and correct major errors.
  • Name the profile and save it.
  • Generate a short test line.
  • Compare engines and regenerate as needed.
  • Export WAV/MP3 when satisfied.

Production checklist

  • Confirm voice consent/rights.
  • Use a clean script with natural punctuation.
  • Generate short sections rather than one giant block.
  • Save good takes immediately.
  • Use effects after the raw voice is good.
  • Preview the full story before export.
  • Label AI-generated voice content when disclosure is appropriate.

Bottom line

Voicebox is presented in the video as a free, local-first AI voice studio that can handle voice profiles, speech generation, story timelines, effects, and hardware acceleration. The best beginner path is simple: install it, record a clean consent-based sample, create a profile, generate a short test, compare engines, then move into stories and effects once the basic voice sounds good.

Best first project: create a 20-second personal voice sample, generate a two-line narration, export it, then repeat with a second engine to compare quality.