Skip to project content
Yulai FanAbout
Product

Voiceover Cutter

I designed and built a tool with the sales team to generate voiceovers or edit recordings against a script, with progress tracking and selective regeneration.

Voiceover Cutter

Overview

Designed and built with sales teammates, saving about 20 minutes per demo video.

Category
Product
Focus
AI-assisted tool + workflow automation
Date
2026

Process

How might we turn a slow, manual voiceover edit into a reusable workflow that preserves finished work and regenerates only what changed?

  1. 01

    Configure the job

    Paste the script, preview its segments, and choose between an uploaded recording or end-to-end AI narration. Voice isolation and narrator roles can be configured before processing.

  2. 02

    Validate and recover

    Check the script and required inputs before creating a job. If an active job already exists, reconnect to it and resume live status updates instead of starting over.

  3. 03

    Process by source

    Synthesize full narration, or transcribe and align a recording to the script, cut matched human clips, and generate the remaining AI voice segments.

  4. 04

    Regenerate selectively

    When only AI voice text changes, reuse the completed human clips and regenerate just the affected narration. Larger edits trigger a new alignment or cut.

  5. 05

    Verify and reuse

    Stitch every segment in script order, verify the output, and provide an activity log and downloadable file. Teammates can finish, retry, edit the script, or create another voiceover.

Voiceover system

Production paths

Paste script
No recording

Generate voices with ElevenLabs

Export M4A

Upload recording

Match recording to script

Remove mistakes and retakes

Insert generated agent voices

Export MP4 or M4A

Script parsing and voice assignment

Script inputProduct behavior
Normal textNarrator voice, or recorded speech in cut mode
Agent 1:ElevenLabs Agent 1 voice
Agent 2:ElevenLabs Agent 2 voice
Unlabeled AI blockAgent 2 fallback voice
Input to output

Normal textNarrator clip

Agent 1:Agent 1 clip

Agent 2:Agent 2 clip

Ordered clipsFinal file

Media processing architecture

Next.js interfaceFastAPI processing serviceElevenLabs + alignment logic + FFmpegGenerated M4A or edited MP4

End-to-end production workflow

How a script moves from source setup through assembly, review, selective regeneration, and export.

End-to-end voiceover production workflowA teammate pastes a script, chooses generated narration or an uploaded recording, applies voice rules to each segment, arranges the clips, reviews the result, and selectively regenerates only the content that changed.Set up sourceInterpret scriptAssembleReview and reuseNo, defaultYes, optionalYesNoNormal textAgent 1Agent 2UnlabeledGenerateCutOpenVoiceover CutterPaste script inoriginal orderUpload arecording?Generate VoicemodeDefault narrator orElevenLabs voice IDCut ExistingSpeech modeVoice isolationenabled?Clean recordingvoice trackUse originalrecordingParse script inoriginal orderSegmenttype?Normal textCurrentmode?Narrator voiceMatch and cut wordsfrom recordingAI VOICE withAgent 1 labelAI VOICE withAgent 2 labelUnlabeledAI VOICE blockElevenLabsAgent 1 voiceElevenLabsAgent 2 voiceAgent 2fallback voiceCreate and arrangeclips in script orderShow live progresswhile generating, cutting,and stitchingSuccessErrorUser stopsDoneCut anotherEdit and re-runYesNoReuse human clipsReuse cleaned mediaProcessingresult?Verify output andactivity detailsDownload M4A,MP4, or M4A cutNextaction?CompleteClear job andstart freshShow error detailsand allow resetCancel at next stageboundary and unlockOnly AI VOICEtext changed?Regenerate onlyagent voicesRe-cut changedhuman text