build and bail (4)

I Was Going to Voice My Kids Podcast Myself. Then I Didn’t.

Two months ago I wrote a post about turning my middle grade novel into a podcast. In it I said the whole thing hinged on cloning my own voice — Mike meets an alternate version of himself every book, so I’d need two Mikes, and an AI clone of me was the obvious answer.

I’m not doing that anymore.

I found a tool called VocalLab.ai on AppSumo, and once I opened the audiobook mode I realized I wasn’t looking at a voice cloning problem. I was looking at a casting problem I could actually afford.

Heads up: the AppSumo link in this post is a referral link. If you buy through it I get a cut. I bought it with my own money before I had any reason to write about it.


What changed

The old plan:

  • Record all 9 episodes myself in Audacity
  • Clone my voice in ElevenLabs for the counterpart character
  • A 43-year-old man doing a 10-year-old’s voice for 2.7 hours

The new plan:

  • Every character gets its own voice
  • I write, direct, and edit — I don’t perform
  • Audacity is now for assembly and timing, not recording

The honest version: I did test recordings. I sound like a 43-year-old man doing a 10-year-old’s voice. Finn Caspian works because that guy can actually do it. I can’t.


How VocalLab’s audiobook mode actually works

This isn’t the standard text-to-speech box. It’s structured for multi-character work:

  • One character block at a time
  • 2,000 character limit per block
  • Inline sound tags dropped right into the text — [laugh], [breathe], [sigh]
  • No steering prompts. The text and the tags are your only direction.

image

That last one matters. If you’ve used TTS tools where you write a separate “read this sadly, slower, more distant” instruction, that doesn’t exist here. Your only lever is the writing itself. Short sentences read slower. Paragraph breaks create air. Punctuation is direction.

Which means the script stopped being a script and became a set of instructions for a machine.


The two things I got wrong first

The pause problem. There’s a moment where Mike says “No. Way.” It has to land as two beats with a real gap. Generated as one block, it comes out as one flat phrase. The fix is three separate blocks — “No.” then silence I insert manually at about 0.8 seconds, then “Way.” Any dramatic beat in your script is a separate generation. Plan for it.

The bleed problem. Grandpa’s laugh was tagged inline at the end of his line and it smeared into the next thing he said. Standalone laughs get their own block with nothing in it but [laugh]. Generate alone, drop into the timeline manually.


The rewrite I didn’t see coming

The book is third person. “Mike opened one eye.” Standard middle grade.

But the podcast has a framing device — a Counterpart Log that opens and closes each episode in Mike’s voice — and that log was already first person. So I had a narrator saying “there’s a journal on my nightstand” and then, thirty seconds later, “Mike chewed his toast.”

It’s a small thing on the page. In audio it’s a guy talking about himself in the third person for 18 minutes.

I went back and forth with Claude on it. Third person is a legitimate audio drama convention — the narrator as an older Mike looking back. Finn Caspian does exactly that. But the log already set the precedent, and first person closes the gap.

Converted all of Episode 1. Mostly a find-and-replace with verb tense cleanup, not a rewrite. Two things I learned doing it:

  • Description blocks stay mixed. “Olivia came stomping in,” not “she came stomping in.” Mike is still observing the scene, and first-person narration does this naturally.
  • The end-of-episode teaser got better. “I learn the rules” beats “Mike learns the rules.” More personal, more hook.

Where it stands

  • Episode 1 is formatted, converted, and partially generated
  • Episodes 2 through 9 are scripted but not yet reformatted for the block structure
  • Still recording all 9 before publishing any of them
  • Still dropping 3 on launch day, weekly after that

Here’s the opening of Episode 1. Judge it yourself.


The stack, updated

  • VocalLab.ai — all character voices (referral link — I get a cut if you buy)
  • Claude — market research, adaptation, scripts, the first-person conversion
  • Audacity — assembly, timing, manual gaps
  • Buzzsprout — hosting and distribution
  • Gemini (Lyria 3) — the music sting

Gone from the last version: ElevenLabs, and me.


The part I’m still uneasy about

I’m making a children’s podcast where no child ever hears a human voice.

I don’t have a clean answer for that. The alternative was a book nobody could find on KDP, or nine episodes of me doing a voice I can’t do. Bad audio doesn’t serve kids either. But I’m not going to pretend this is a neutral tradeoff, and I’m not going to hide it in the credits.

Tell me I’m wrong in the comments. I’ll read them.

Recording continues. Unless I bail.

Leave a Comment

Your email address will not be published. Required fields are marked *