AI voiceover quality assurance

Compare an AI voiceover with its source script

Use Voiceover QA to find missing, repeated, or changed words plus audio anomalies at exact timestamps before regenerating or publishing. English comparison is production-ready; Japanese post-generation comparison remains beta.

Play the original eight-second synthetic demo audio. It contains one known omitted word, one 2.45-second pause, and one changed word; it contains no customer work or testimonial.

Core workflow and supported input

  1. Sign in, paste one source script, and upload the matching single-speaker narration.
  2. Use MP3, WAV, or M4A up to 100 MB and 30 minutes.
  3. The private server queue transcribes the audio, aligns detected words, and checks decoded audio signals.
  4. Review, confirm, or ignore each candidate and export the timestamped repair list as CSV.

Output

The report can surface missing, extra, repeated, and changed wording plus long silence, clipping, and abrupt level changes. Each item includes a timestamp, source context, detected text when available, confidence, and a suggested review action.

Privacy and retention boundary

Uploaded audio is private, stored outside the public web root, sent server-to-server to the configured speech-to-text provider, and scheduled for deletion after about 24 hours. Analytics does not receive audio, script or transcript text, filenames, project or customer names, payment identifiers, or finding details.

Limits and recovery

Findings are candidates for human review. Provider transcripts may normalize wording or omit repeated tokens. Voiceover QA does not judge acting, emotion, native pronunciation, naturalness, publication readiness, client approval, or commercial results. Empty, unsupported, oversized, provider-failed, timed-out, insufficient-credit, and failed-payment states provide a retry or support path.

Read the recovery guide, model boundaries, authoritative sources, and current billing behavior.