Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

2 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

oc2 promo txt_img

Script to Voice Generator — Hume AI Octave 2 TTS

Welcome to any script writing author out there!

Convert formatted script files into fully voiced audio using Hume AI Octave 2 TTS — a high-quality cloud text-to-speech engine with expressive, emotionally nuanced voice generation.

Get it here: https://reactorcore.itch.io/script-to-voice-generator-hume-ai

Made by Reactorcorehttps://linktr.ee/reactorcore


stvg humeai screenshot (1) stvg humeai screenshot (2)

What It Does

Script to Voice Generator reads a formatted .txt or .md script file and:

  • Converts each dialogue line to speech using Hume AI Octave 2 (any voice from the Hume AI library or your custom voices).
  • Saves individual clips for each line — both clean (TTS only) and effects-processed.
  • Merges all clips into a single audio file, with smart pauses based on punctuation.
  • Produces both a raw merge and a loudness-normalized merge.
  • Generates a reference sheet listing every clip filename and its spoken text.

Multiple speakers are supported. Each speaker gets their own voice, acting description, pitch, speed, and audio effects settings, stored in character_profiles.json so they're remembered between sessions.


System Requirements

  • Windows 11 — built and tested on Windows 11.
  • Windows 10 — untested, use at your own risk.
  • Linux / macOS — the compiled .exe is Windows-only. No Linux or macOS build is available.
  • Internet connection — required for TTS generation (Hume AI is a cloud API).

What You Need

Hume AI API key — Required for TTS generation.

  1. Go to app.hume.ai/developers and create an account.
  2. Click Create API Key, give it a name, and copy it.
  3. Open the Settings tab (Tab 4) in the app and paste it in.

Only the API Key is needed — not the Secret Key (that's for OAuth token exchange, not needed here).

FFMPEG — Required for audio effects and merging.

Python 3.x — Required to run from source (not needed if using the compiled .exe). Use build_exe.bat to build to a single .exe in one click.

A note on Hume AI pricing

Hume AI is a cloud API service — you pay per character of text generated. The free tier gives roughly 10,000 characters per month (~10 minutes of audio). Paid tiers start at $3/month (30,000 chars) and scale from there.

Octave 2 produces emotionally nuanced, cinematic voice that's difficult to achieve with local models. The recommendation is to use it intentionally — write tight scripts, use acting descriptions to shape delivery, and test individual voices before committing to a full run.

Check Hume AI pricing and usage at: https://app.hume.ai/billing

Caution: Test Voice clips in Tab 2 also consume your character quota. Use them wisely.


Quick Start

1. Set your API key (Tab 4)

Launch the app and go to Tab 4 — Settings. Paste your API key in the Hume AI API section and click Save Key. Voices will load automatically.

2. Write or prepare a script file

Scripts are .txt or .md files. Each spoken line uses the format:

SpeakerID: Dialogue text goes here.

Example:

# My Short Film

Alex: Hey, are you okay?
Jordan: Yeah, I'm fine. [sighs] Just tired.
(1.0s)
Alex: You sure? You look pale.
Jordan: [barely suppressing frustration] I said I'm fine.

See Script Format below for full syntax details.

3. Load the script (Tab 1)

  1. Launch the program and click Open Script File.
  2. The parser checks for formatting errors and lists them in the log.
  3. Fix any errors in your text editor and click Reload Script.
  4. When the parse log shows no errors, click Continue →.

4. Configure voices (Tab 2)

Each detected speaker gets a panel with:

  • Voice — Choose from any voice in the Hume AI library or your custom voices. Voices are fetched automatically when the app loads with a valid API key.
  • Acting Description — A short phrase describing how this speaker talks (e.g. "Warm, conspiratorial, slightly amused"). Applies to all their lines. Override per-line in the script using [brackets]. (Currently not transmitted to Hume AI — Octave 2 is in preview and rejects this field. Stored for when support is restored.)
  • Speed — Speaking rate from 0.75× to 1.5×.
  • Pitch — Multiplier from ×0.5 to ×2.0 (FFMPEG rubberband pitch shift). Default ×1.0 = no shift.
  • Level — 5–100% relative volume. 100% = full normalized output (default). Reduce to make a speaker quieter in the mix.
  • Yell Impact — Slows down single-word exclamatory lines (e.g. YES!). Makes such lines sound more deliberate and impactful. Set per speaker.
  • Audio Effects — Radio, Reverb, Distortion, Telephone, Robot Voice, Cheap Mic, Underwater, Megaphone, Worn Tape, Intercom, Alien Voice, Cave, and Pitch Shift. Most effects have Off / Mild / Medium / Strong levels.

Use Test Voice to generate a quick preview clip and hear the settings immediately.

Settings auto-save to character_profiles.json on every change, so known speakers are recalled automatically next session.

5. Generate (Tab 3)

  1. Enter a Project Name (used as a filename prefix, 20 chars max).
  2. Choose an Output Folder.
  3. Click Generate All and confirm.

The generation log shows progress. When done, all files appear in the output folder:

output_folder/
├── clips_clean/          ← Raw TTS clips (no FFMPEG effects)
│   └── project_0001_Speaker_line-text.mp3
├── clips_effect/         ← Effects-processed clips (peak-normalized)
│   └── project_0001_Speaker_line-text.mp3
├── sfx/                  ← Processed SFX copies (only if SFX effects active)
├── !project_merged_pure.mp3        ← Merged audio, no normalization
├── !project_merged_loudnorm.mp3    ← Merged audio, loudness-normalized
└── project_reference.txt           ← Line-by-line reference sheet

Script Format

Dialogue lines

SpeakerID: Spoken text goes here.
  • SpeakerID must be 20 characters or fewer. Allowed: letters, numbers, spaces, hyphens, underscores.
  • All text after the first colon is spoken. Additional colons in the line are fine.
  • Lines over 4000 characters throw a parse error.

Acting descriptions (brackets)

Alex: I can't believe this. [sighs heavily]
Jordan: [whispering] Don't move.
Narrator: [urgent, barely breathing] It was already too late.

[brackets] on a dialogue line are extracted as acting direction. They are not spoken as words and do not appear in the reference sheet text.

Multiple bracket groups on one line are joined automatically:

Narrator: [tense] [leaning in close] Keep your voice down.
→ description: "tense, leaning in close"
→ spoken text: "Keep your voice down."

The Tab 2 Acting Description field sets a default for all lines from a speaker. A [bracket] in the script overrides it for that specific line only.

Octave 2 preview limitation: The description field is currently rejected by the Hume AI API. Acting descriptions are parsed and stored correctly, but are not transmitted during generation. This is expected to be resolved as Octave 2 matures out of preview.

Tips for effective descriptions (for when this is restored):

  • Specific beats generic: "melancholy" beats "sad", "barely suppressing a smirk" beats "amused"
  • Describe emotional state and context, not pace — use the speed slider for pace
  • Keep under 100 characters total; 1–2 descriptors per line is plenty
  • Combining context works well: [urgent whisper — someone might be listening], [speaking to a child]

Hume in-text pause tags

Jordan: I've been thinking about this for a long time. [pause] A very long time.

[pause] and [long pause] are Hume delivery hints sent inside the spoken text. They are not extracted as acting description — they remain in the text as-is. They're distinct from STVG's (Xs) pause syntax, which inserts silence in the merged audio at the merger level.

Headings

# Scene title
## Sub-scene

Treated as metadata. Sets the script title. Not voiced.

Comments

// This is a comment
/* Multi-line
   comment */

Not voiced. Useful for stage directions, notes, or commented-out lines.

Pauses

(1.5s)
(pause 2.0)
(0.8)

Any line that is only parentheses containing a number inserts a silent pause in the merged audio. The number is in seconds.

Sound effects

{play filename.mp3, c1, loop}
{stop c1}
{stop all}
{play explosion.wav, c2, once}

Sound effect events are placed in the merge timeline at the correct position. Sound effect files must exist in the SFX folder specified in Tab 2.

Note: If a sound effect is the very last item in your script, it needs a pause after it to actually be heard in the merged audio. Add a (pause) line equal to or longer than the sound effect's duration immediately after the {play} line. Without it, the base audio ends at the same moment the SFX starts, and the SFX gets cut off.

Like this:

Rei: Signing off.
{play cloth.wav, c1, once}
(2.0s)

Supported formats — Any audio format FFMPEG can read: .mp3, .wav, .ogg, .flac, .aac, .m4a, and others. The filename in your script must match the actual file exactly (including extension).

Inner thoughts

SpeakerID: (( This line is an inner thought. ))

Wrapping dialogue in double parentheses marks it as an inner thought. Inner thought lines are voiced with a special filtering effect configured in Tab 4 (Dissociated, Whisper, or Dreamlike presets, or Custom). The filter runs on top of all the speaker's regular effects.

Inline notation

  • [brackets] on a dialogue line are acting direction — extracted and stored, not spoken, not shown in reference text. Not currently transmitted to Hume AI (Octave 2 preview limitation). Exception: [pause] / [long pause] are delivery tags that pass through as text.
  • **bold**, _italic_, and ~~strikethrough~~ markers are stripped before TTS.
  • // after dialogue text starts an inline comment; everything after it is stripped. A space before // is required (so URLs are not accidentally stripped).

Settings Tab (Tab 4)

Hume AI API — Enter your API key here. Voices load automatically once a valid key is saved. Use "Open Billing Page" to check your usage on the Hume AI dashboard.

Generation Options — Clip continuity chains each clip's generation_id for more natural voice flow between consecutive lines from the same speaker.

Silence Trim — Controls how leading/trailing silence is removed from each TTS clip. Default: trim beginning and end. Options: Off, Beginning only, End only, Beginning + End, All silence.

Merged Audio Pauses — Adjust the pause duration added after each punctuation type (period, comma, exclamation, question, hyphen, ellipsis, etc.).

Contextual Modifiers — Fine-tune how pause lengths are modified by context: speaker changes, short lines, long lines, inner thought padding, same-speaker reduction, first/last line padding.

Inner Thoughts Effect — Choose from Whisper, Dreamlike, Dissociated presets or configure custom highpass/lowpass/echo parameters for the inner thought audio filter.


Audio Effects Reference

Effect Description
Radio Filter Walkie-talkie / comms radio effect. Bandpass + phaser + compression.
Reverb Spatial depth. Configurable echo chains.
Distortion Aggressive, gritty clipping and bit crushing.
Telephone Lo-fi compressed sound. Narrow bandpass + bit crushing.
Robot Voice Ring modulator for mechanical / robotic character.
Cheap Mic Degraded quality, poor recording simulation.
Underwater Muffled, wet, submerged sound. Lowpass + flanger.
Megaphone Projected bullhorn. Treble-boosted, punchy, bandpassed.
Worn Tape VHS/cassette degradation. Wow-flutter, lo-fi analog warble.
Intercom Hallway speaker box. Flat, compressed, confined. Adds crackling static noise.
Alien Voice Non-human vocal quality. Three variants: Insectoid, Dimensional, Warble.
Cave Physical stone space reverb. Three variants: Tunnel, Cave, Abyss.
Pitch Shift FFMPEG rubberband pitch shift. Multiplier ×0.5–×2.0. Works independently of speed.

Most effects have Off / Mild / Medium / Strong presets. Alien and Cave use named variants instead. Effects are combinable.


Tips

  • 160+ voices available — The Hume AI library ships with a large built-in voice set. Use the Test Voice button to audition before committing to a full run.

  • Acting descriptions (currently limited) — The description field is rejected by the Octave 2 API in its current preview state. [brackets] and Tab 2 descriptions are still parsed and stored — they'll transmit automatically once Hume restores support.

  • No stability or similarity sliders — Hume Octave 2 doesn't have these parameters. Expressiveness is shaped through speed and punctuation.

  • Clip continuity — Leave "Use clip continuity" on in Tab 4. It chains each generated clip's generation_id to the next, producing more natural voice flow across consecutive lines from the same speaker.

  • Pitch for pitch shifting — The pitch slider in Tab 2 uses FFMPEG rubberband pitch shifting. It is independent of speaking speed.

  • Test each voice before generating everything. The Test Voice button in Tab 2 saves a preview clip and opens it immediately.

  • Cheap Mic at Mild is a subtle effect that adds a hint of realism. Worth trying as a default.

  • Prompt templates — The !docs/prompt_templates/ folder has templates for using AI chatbots to write scripts or generate voice line banks. Open them in any text editor.


Included Docs (!docs/)

Guides

File Contents
!docs/guides/Script_Writing_Guide.md Writing for TTS, pacing with punctuation and pauses, acting descriptions, Hume pause tags, using effects as character design, AI-assisted workflow
!docs/guides/Audio_Effects_Guide.md Full reference for all effects, preset levels, FFMPEG pipeline, Yell Impact, troubleshooting

Example Scripts

Ready-to-load .md script files — open any of them in Tab 1 to see the format in action.

File What it demonstrates
example_tiny.md Minimal 2-line script
example_small.md Short 2-character scene with SFX, pause, and comments
example_full_drama.md Full multi-character drama with SFX channels, inner thoughts, and scene structure
example_monologue.md Single narrator, no character interaction
example_meditation.md Atmospheric piece with long pauses and inner thought lines
example_oneliners.md Voice bank format — one character, many independent lines by category
example_game_scenes.md Multi-scene game dialogue with tactical characters, SFX, and inner thoughts

Prompt Templates

Fill-in-the-blank prompts for generating scripts with an AI chatbot. Copy, fill in characters/scenario, paste to a chatbot, save the output as a .md file, load in Tab 1.

File Use case
cohesive_script.md Continuous scene — characters talk to each other
separate_voice_lines.md Voice bank — independent lines per category
game_scene_pack.md Single game scene with character roles, SFX, and inner thoughts
narrator_monologue.md Single narrator — story, documentary, speech, essay
podcast_interview.md Two-person host/guest conversation
ambient_narration.md Slow, atmospheric, mood-driven spoken word

Troubleshooting

No voices in Tab 2 — Check that your API key is entered and saved in Tab 4. The app fetches voices from Hume AI on startup. If the key is invalid or missing, the voice list will be empty.

FFMPEG not found — Install FFMPEG and make sure it is in your system PATH. Use the automatic installer at https://reactorcore.itch.io/ffmpeg-to-path-installer then restart the program.

Parse errors on load — The parse log in Tab 1 lists every error with line numbers. Fix them in your text editor and click Reload Script.

Voice too quiet — The post-effects normalization pass ensures consistent loudness. If a speaker still sounds quiet relative to others, their Level slider may be below 100%.

Missing voice lines in output — Check the generation log in Tab 3 for per-line errors. An API error or FFMPEG issue on a specific line will be noted.

Test Voice not opening — The file is saved to output_test/ in the program folder. Open it manually if the auto-open fails.

Generation fails / API errors — Check your usage at app.hume.ai/billing to see if you've hit your monthly character quota. Check that your API key is correct and active.


Credits

  • Hume AI Octave 2 TTS — Cloud TTS engine
  • ttkbootstrap — Modern themed tkinter UI
  • FFMPEG — Audio processing and merging
  • Script to Voice Generator — By Reactorcore

Links

Check out everything else I do: https://linktr.ee/reactorcore

About

Convert scripts into fully voiced audio using Hume AI Octave 2 TTS with expressive, emotion-aware voices.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages