ffmpegaudiovideo-automation

Replace audio in video by API: the mux is the easy part

··10 min read
Replace audio in video by API: the mux is the easy part

ElevenLabs shipped Eleven v4 on September 28, 2026, on the free tier, so a lot of builders are regenerating narration for videos they already published. Generating the new tracks is the fast part. What eats the afternoon is that the new WAV is never the same length as the picture, and the captions are already baked into the pixels against the old audio.

Quick answer: A replace audio in video API does three things: strip the original audio stream, mux the new track in, and reconcile the duration difference between the two. The FFmpeg equivalent is ffmpeg -i in.mp4 -i vo.wav -map 0:v:0 -map 1:a:0 -c:v copy -c:a aac -b:a 192k -shortest out.mp4, where -shortest quietly truncates whichever stream runs long. Measure both durations with ffprobe before you mux so you pad or trim on purpose, and send the job to FFmpeg Micro as one call if you'd rather not host an encoder or babysit a batch of long renders.

Why a regenerated voiceover breaks an already-rendered video

Conventional wisdom says replacing audio is a stream-selection trick: point -map at the video from file one and the audio from file two, copy the video, done. That's correct, and for a 1:1 swap of equal-length files it's the whole answer. It's also where every page on this topic stops.

The failure isn't in the mux. It's that text-to-speech output has its own pacing. v4 ships inline performance direction tags and voice cloning from about 10 seconds of reference audio, per ElevenLabs' launch post, and a delivery that pauses for effect runs long. Regenerate a 90-second script and you routinely get 83 seconds or 98 seconds back. That's a 5 to 10 percent delta, not rounding.

Picture a faceless history channel with 180 rendered MP4s, each with a flat 2025-era narration track and hard-burned captions. With v4 Turbo at roughly 150 ms median time to first speech, regenerating all 180 scripts finishes over lunch. Re-rendering the 180 videos is the week. The real job is a decision tree over durations plus a caption rebuild, and neither has anything to do with -map.

Measure the drift before you mux

Two ffprobe calls give you everything the decision needs. They read container headers rather than decoding, so they return in milliseconds on a 400 MB file.

V=$(ffprobe -v error -show_entries format=duration -of csv=p=0 video.mp4)
A=$(ffprobe -v error -show_entries format=duration -of csv=p=0 vo.wav)
echo "video=$V audio=$A delta=$(echo "$A - $V" | bc) ratio=$(echo "scale=4; $A / $V" | bc)"

The delta decides whether you pad or trim, and the exact video duration is what you pass to apad. Guessing a pad value by hand is the most common reason a batch ships with four seconds of dead air on every clip.

The drift decision: pad, trim, re-cut, and when atempo is the wrong answer

The size of the delta tells you which of four actions is honest, and only the first two are a mux. Every FFmpeg guide to length mismatch recommends an atempo factor of audio length over video length. The ffmpeg-cookbook version uses atempo=1.001 to fit 100.1 seconds of audio into 100 seconds of video, and 0.1 percent is inaudible. Over-generalized to the 8 percent deltas TTS produces, it's a chipmunk. The filter also only accepts 0.5 to 2.0 and has to be chained beyond that range.

DeltaUsual causeAction
Under 0.5 sTrailing silence, encoder rounding`apad` to the exact video duration, or `atempo`
0.5 s to 3 sNew delivery breathes differentlyPad with silence, or trim the tail after checking it's silent
3 s to 10%Script wording changedExtend or trim the picture, not the audio
Over 10%Different scriptRe-cut the video; this is an edit, not a mux

When the new voiceover is shorter, pad it to the picture and keep the video stream untouched:

ffmpeg -i video.mp4 -i vo.wav \
  -filter_complex "[1:a]apad=whole_dur=92.6[a]" \
  -map 0:v:0 -map "[a]" -c:v copy -c:a aac -b:a 192k out.mp4

whole_dur takes the target total length, the $V you just probed. pad_dur takes the amount of silence to add instead. Either way the explicit -map of the filtered stream matters: with a filter graph in play, FFmpeg no longer picks streams for you.

When the new voiceover is longer, you have to add picture. Holding the final frame is the cheapest option that doesn't lie about the content:

ffmpeg -i video.mp4 -i vo.wav \
  -filter_complex "[0:v]tpad=stop_mode=clone:stop_duration=4.2[v]" \
  -map "[v]" -map 1:a:0 -c:v libx264 -crf 20 -preset veryfast \
  -c:a aac -b:a 192k -shortest out.mp4

That one costs a full video re-encode, because filtering and stream copy can't be combined on the same stream. Some hosted replace-audio endpoints make this decision without exposing it: longer audio gets trimmed, shorter audio leaves silence, no parameter. A narration that runs long loses its last two sentences and nothing in the response tells you.

Rebuild the captions from the TTS timings, not a second transcription

Burned-in captions are pixels in the picture, so there's no offset to shift and no track to re-time. Matching captions to new audio means a fresh subtitle file and a re-export from a text-free master. If your only copy is the captioned render, the clip goes back through assembly.

The part people pay for twice is the transcription, and you don't need it. The call that generated the audio hands you the timings: ElevenLabs' /v1/text-to-speech/{voice_id}/with-timestamps returns an alignment object containing characters, character_start_times_seconds, and character_end_times_seconds alongside the audio. For tracks you already rendered without timestamps, the Forced Alignment API takes the audio plus your script and returns word- and character-level timings with an alignment loss score. Either way the caption rebuild is a data transform, not another Whisper bill.

Character arrays into SRT cues, as an n8n Code node:

const { characters, character_start_times_seconds: starts,
        character_end_times_seconds: ends } = $json.alignment;

const words = [];
let cur = null;
characters.forEach((ch, i) => {
  if (/\s/.test(ch)) { cur = null; return; }
  if (!cur) { cur = { text: '', start: starts[i], end: ends[i] }; words.push(cur); }
  cur.text += ch;
  cur.end = ends[i];
});

const MAX = 6;
const pad = (s) => new Date(s * 1000).toISOString().substr(11, 12).replace('.', ',');
const srt = [];
for (let i = 0; i < words.length; i += MAX) {
  const g = words.slice(i, i + MAX);
  srt.push(`${srt.length + 1}\n${pad(g[0].start)} --> ${pad(g[g.length - 1].end)}\n` +
           g.map(w => w.text).join(' ') + '\n');
}
return [{ json: { srt: srt.join('\n') } }];

Then burn it onto the swapped file. Font styling is where this silently fails in containers: subtitles and drawtext need fontconfig and a font package installed, or you get a clean exit and no visible text.

ffmpeg -i swapped.mp4 -vf "subtitles=new.srt:force_style='FontName=DejaVu Sans,FontSize=26,Outline=2'" \
  -c:v libx264 -crf 20 -c:a copy captioned.mp4

Swap video audio in n8n or Make with HTTP calls

n8n and Make can't shell out to FFmpeg on a hosted plan. n8n disabled the Execute Command node by default in v2.0, and Make and Zapier never had it, so the pipeline is HTTP nodes end to end. The sequence for a back catalogue:

  1. Read the row (clean master URL, script text, voice ID) from Sheets, Airtable, or Postgres.
  2. POST the script to ElevenLabs with-timestamps, store the audio to S3 or Drive, keep the alignment object in the item.
  3. Probe both durations and compute delta and ratio in a Code node.
  4. Branch with an If node on the thresholds above. Pad, trim, or route long-delta rows to a human review queue.
  5. Submit the mux job, then the caption burn, as HTTP requests against a video API and wait on the webhook.
  6. Write the output URL back to the row so the batch is resumable.

Step 5 is where a hosted job API earns its place. FFmpeg Micro takes the input URLs and the filter work as one call per job and returns a webhook when the output is ready, so there's no encoder to install, no 20-minute execution to keep alive, and the MP4 never enters n8n as binary data. That's what kills 180-file batches: n8n holds binaries in memory, the same reason large video transcription fails when you send the file instead of a URL.

Common pitfalls

The music bed surprises people most. If the original mix had music or room tone under the narration, replacing the whole audio stream throws it away. Mix the old track down instead: [0:a]volume=0.12[bed];[1:a][bed]amix=inputs=2:duration=first:dropout_transition=0[a]. With the voiceover as the first input, duration=first ends the mix on the narration.

Double re-encoding is the second. Pad, mux, then burn captions as three separate commands and the video gets encoded twice. Build one filter graph with tpad and subtitles, or accept the generation loss and raise your CRF budget.

Sample rate is the quiet one. TTS output is often 44.1 kHz while your library is 48 kHz, and some players handle the mismatch badly. Add -ar 48000 to the audio encode.

Last, check the tail before you trim. A -t cut assumes the extra audio is silence. If v4 added a closing sentence your 2025 script didn't have, trimming deletes it and the output still passes every automated check.

When replacing the audio is the wrong move

An audio swap is only valid when the picture doesn't depend on the old audio. It isn't valid for an on-camera talking head, where the mouth is synced to the track you're removing; for a video cut to the beats of the old narration, where b-roll changes land on sentences that no longer exist; or for a delta past 10 percent, where the script changed and the clip needs re-cutting. For those, the right tool is a composition step that rebuilds the timeline from source assets, not a mux.

FAQ

Can I just speed up the new voiceover so it fits the video?

Speeding up the voiceover works below about 1 or 2 percent and nowhere above it. atempo is the right fix for sub-percent drift, but at the 5 to 20 percent deltas a regenerated TTS track produces, it flattens the performance direction you regenerated for.

Do I have to re-transcribe the new audio to fix the captions?

Re-transcription is unnecessary. Request the audio from ElevenLabs via the with-timestamps endpoint and the character-level start and end times come back with it, or run Forced Alignment against an already-rendered track using your existing script.

How do I replace video audio in n8n without a community node?

Replacing video audio in n8n needs only the HTTP Request and Code nodes: POST the job to a video API with the video and audio URLs, then wait on the webhook. The Execute Command node has been disabled by default since n8n v2.0, and hosted n8n can't run FFmpeg.

Will swapping the audio re-encode the video?

A plain audio swap keeps -c:v copy, so the video stream is untouched and the job finishes in seconds. Adding any video filter, tpad to hold a frame or subtitles to burn captions, forces a full re-encode.

What if the original video has music I want to keep?

Keep the music by mixing instead of replacing: lower the original stream with volume and combine it with the new narration through amix, rather than mapping only the new track. A music bed can't be separated from narration inside an already-mixed stream, so if the music must stay clean you need the original stems.

Probe, decide, pad, mux, re-burn: five steps, all of them HTTP calls you can fire from a workflow node or an AI agent tool call. The free tier covers running one back-catalogue video end to end before you point it at the other 179.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Ready to process videos at scale?

Start using FFmpeg Micro's simple API today. No infrastructure required.

Get Started Free