Replace audio in video by API: the mux is the easy part

ElevenLabs shipped Eleven v4 on September 28, 2026, on the free tier, so a lot of builders are regenerating narration for videos they already published. Generating the new tracks is the fast part. What eats the afternoon is that the new WAV is never the same length as the picture, and the captions are already baked into the pixels against the old audio.
Quick answer: A replace audio in video API does three things: strip the original audio stream, mux the new track in, and reconcile the duration difference between the two. The FFmpeg equivalent isffmpeg -i in.mp4 -i vo.wav -map 0:v:0 -map 1:a:0 -c:v copy -c:a aac -b:a 192k -shortest out.mp4, where-shortestquietly truncates whichever stream runs long. Measure both durations withffprobebefore you mux so you pad or trim on purpose, and send the job to FFmpeg Micro as one call if you'd rather not host an encoder or babysit a batch of long renders.
Why a regenerated voiceover breaks an already-rendered video
Conventional wisdom says replacing audio is a stream-selection trick: point -map at the video from file one and the audio from file two, copy the video, done. That's correct, and for a 1:1 swap of equal-length files it's the whole answer. It's also where every page on this topic stops.
The failure isn't in the mux. It's that text-to-speech output has its own pacing. v4 ships inline performance direction tags and voice cloning from about 10 seconds of reference audio, per ElevenLabs' launch post, and a delivery that pauses for effect runs long. Regenerate a 90-second script and you routinely get 83 seconds or 98 seconds back. That's a 5 to 10 percent delta, not rounding.
Picture a faceless history channel with 180 rendered MP4s, each with a flat 2025-era narration track and hard-burned captions. With v4 Turbo at roughly 150 ms median time to first speech, regenerating all 180 scripts finishes over lunch. Re-rendering the 180 videos is the week. The real job is a decision tree over durations plus a caption rebuild, and neither has anything to do with -map.
Measure the drift before you mux
Two ffprobe calls give you everything the decision needs. They read container headers rather than decoding, so they return in milliseconds on a 400 MB file.
V=$(ffprobe -v error -show_entries format=duration -of csv=p=0 video.mp4)
A=$(ffprobe -v error -show_entries format=duration -of csv=p=0 vo.wav)
echo "video=$V audio=$A delta=$(echo "$A - $V" | bc) ratio=$(echo "scale=4; $A / $V" | bc)"
The delta decides whether you pad or trim, and the exact video duration is what you pass to apad. Guessing a pad value by hand is the most common reason a batch ships with four seconds of dead air on every clip.
The drift decision: pad, trim, re-cut, and when atempo is the wrong answer
The size of the delta tells you which of four actions is honest, and only the first two are a mux. Every FFmpeg guide to length mismatch recommends an atempo factor of audio length over video length. The ffmpeg-cookbook version uses atempo=1.001 to fit 100.1 seconds of audio into 100 seconds of video, and 0.1 percent is inaudible. Over-generalized to the 8 percent deltas TTS produces, it's a chipmunk. The filter also only accepts 0.5 to 2.0 and has to be chained beyond that range.
| Delta | Usual cause | Action |
|---|---|---|
| Under 0.5 s | Trailing silence, encoder rounding | `apad` to the exact video duration, or `atempo` |
| 0.5 s to 3 s | New delivery breathes differently | Pad with silence, or trim the tail after checking it's silent |
| 3 s to 10% | Script wording changed | Extend or trim the picture, not the audio |
| Over 10% | Different script | Re-cut the video; this is an edit, not a mux |
When the new voiceover is shorter, pad it to the picture and keep the video stream untouched:
ffmpeg -i video.mp4 -i vo.wav \
-filter_complex "[1:a]apad=whole_dur=92.6[a]" \
-map 0:v:0 -map "[a]" -c:v copy -c:a aac -b:a 192k out.mp4
whole_dur takes the target total length, the $V you just probed. pad_dur takes the amount of silence to add instead. Either way the explicit -map of the filtered stream matters: with a filter graph in play, FFmpeg no longer picks streams for you.
When the new voiceover is longer, you have to add picture. Holding the final frame is the cheapest option that doesn't lie about the content:
ffmpeg -i video.mp4 -i vo.wav \
-filter_complex "[0:v]tpad=stop_mode=clone:stop_duration=4.2[v]" \
-map "[v]" -map 1:a:0 -c:v libx264 -crf 20 -preset veryfast \
-c:a aac -b:a 192k -shortest out.mp4
That one costs a full video re-encode, because filtering and stream copy can't be combined on the same stream. Some hosted replace-audio endpoints make this decision without exposing it: longer audio gets trimmed, shorter audio leaves silence, no parameter. A narration that runs long loses its last two sentences and nothing in the response tells you.
Rebuild the captions from the TTS timings, not a second transcription
Burned-in captions are pixels in the picture, so there's no offset to shift and no track to re-time. Matching captions to new audio means a fresh subtitle file and a re-export from a text-free master. If your only copy is the captioned render, the clip goes back through assembly.
The part people pay for twice is the transcription, and you don't need it. The call that generated the audio hands you the timings: ElevenLabs' /v1/text-to-speech/{voice_id}/with-timestamps returns an alignment object containing characters, character_start_times_seconds, and character_end_times_seconds alongside the audio. For tracks you already rendered without timestamps, the Forced Alignment API takes the audio plus your script and returns word- and character-level timings with an alignment loss score. Either way the caption rebuild is a data transform, not another Whisper bill.
Character arrays into SRT cues, as an n8n Code node:
const { characters, character_start_times_seconds: starts,
character_end_times_seconds: ends } = $json.alignment;
const words = [];
let cur = null;
characters.forEach((ch, i) => {
if (/\s/.test(ch)) { cur = null; return; }
if (!cur) { cur = { text: '', start: starts[i], end: ends[i] }; words.push(cur); }
cur.text += ch;
cur.end = ends[i];
});
const MAX = 6;
const pad = (s) => new Date(s * 1000).toISOString().substr(11, 12).replace('.', ',');
const srt = [];
for (let i = 0; i < words.length; i += MAX) {
const g = words.slice(i, i + MAX);
srt.push(`${srt.length + 1}\n${pad(g[0].start)} --> ${pad(g[g.length - 1].end)}\n` +
g.map(w => w.text).join(' ') + '\n');
}
return [{ json: { srt: srt.join('\n') } }];
Then burn it onto the swapped file. Font styling is where this silently fails in containers: subtitles and drawtext need fontconfig and a font package installed, or you get a clean exit and no visible text.
ffmpeg -i swapped.mp4 -vf "subtitles=new.srt:force_style='FontName=DejaVu Sans,FontSize=26,Outline=2'" \
-c:v libx264 -crf 20 -c:a copy captioned.mp4
Swap video audio in n8n or Make with HTTP calls
n8n and Make can't shell out to FFmpeg on a hosted plan. n8n disabled the Execute Command node by default in v2.0, and Make and Zapier never had it, so the pipeline is HTTP nodes end to end. The sequence for a back catalogue:
- Read the row (clean master URL, script text, voice ID) from Sheets, Airtable, or Postgres.
- POST the script to ElevenLabs
with-timestamps, store the audio to S3 or Drive, keep the alignment object in the item. - Probe both durations and compute
deltaandratioin a Code node. - Branch with an If node on the thresholds above. Pad, trim, or route long-delta rows to a human review queue.
- Submit the mux job, then the caption burn, as HTTP requests against a video API and wait on the webhook.
- Write the output URL back to the row so the batch is resumable.
Step 5 is where a hosted job API earns its place. FFmpeg Micro takes the input URLs and the filter work as one call per job and returns a webhook when the output is ready, so there's no encoder to install, no 20-minute execution to keep alive, and the MP4 never enters n8n as binary data. That's what kills 180-file batches: n8n holds binaries in memory, the same reason large video transcription fails when you send the file instead of a URL.
Common pitfalls
The music bed surprises people most. If the original mix had music or room tone under the narration, replacing the whole audio stream throws it away. Mix the old track down instead: [0:a]volume=0.12[bed];[1:a][bed]amix=inputs=2:duration=first:dropout_transition=0[a]. With the voiceover as the first input, duration=first ends the mix on the narration.
Double re-encoding is the second. Pad, mux, then burn captions as three separate commands and the video gets encoded twice. Build one filter graph with tpad and subtitles, or accept the generation loss and raise your CRF budget.
Sample rate is the quiet one. TTS output is often 44.1 kHz while your library is 48 kHz, and some players handle the mismatch badly. Add -ar 48000 to the audio encode.
Last, check the tail before you trim. A -t cut assumes the extra audio is silence. If v4 added a closing sentence your 2025 script didn't have, trimming deletes it and the output still passes every automated check.
When replacing the audio is the wrong move
An audio swap is only valid when the picture doesn't depend on the old audio. It isn't valid for an on-camera talking head, where the mouth is synced to the track you're removing; for a video cut to the beats of the old narration, where b-roll changes land on sentences that no longer exist; or for a delta past 10 percent, where the script changed and the clip needs re-cutting. For those, the right tool is a composition step that rebuilds the timeline from source assets, not a mux.
FAQ
Can I just speed up the new voiceover so it fits the video?
Speeding up the voiceover works below about 1 or 2 percent and nowhere above it. atempo is the right fix for sub-percent drift, but at the 5 to 20 percent deltas a regenerated TTS track produces, it flattens the performance direction you regenerated for.
Do I have to re-transcribe the new audio to fix the captions?
Re-transcription is unnecessary. Request the audio from ElevenLabs via the with-timestamps endpoint and the character-level start and end times come back with it, or run Forced Alignment against an already-rendered track using your existing script.
How do I replace video audio in n8n without a community node?
Replacing video audio in n8n needs only the HTTP Request and Code nodes: POST the job to a video API with the video and audio URLs, then wait on the webhook. The Execute Command node has been disabled by default since n8n v2.0, and hosted n8n can't run FFmpeg.
Will swapping the audio re-encode the video?
A plain audio swap keeps -c:v copy, so the video stream is untouched and the job finishes in seconds. Adding any video filter, tpad to hold a frame or subtitles to burn captions, forces a full re-encode.
What if the original video has music I want to keep?
Keep the music by mixing instead of replacing: lower the original stream with volume and combine it with the new narration through amix, rather than mapping only the new track. A music bed can't be separated from narration inside an already-mixed stream, so if the music must stay clean you need the original stems.
Probe, decide, pad, mux, re-burn: five steps, all of them HTTP calls you can fire from a workflow node or an AI agent tool call. The free tier covers running one back-catalogue video end to end before you point it at the other 179.
About Javid Jamae
Founder & CEO at FFmpeg Micro
Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.
You might also like

16kHz Mono WAV for Whisper, Without a Local FFmpeg Binary
The FFmpeg 16kHz mono WAV command Whisper needs, plus the pipeline half: pulling audio from an MP4, size math, channel picks, and no local binary to install.

Extract the Last Frame in FFmpeg to Chain Veo 3 Clips Cleanly
Extract last frame FFmpeg commands return the wrong frame after a fast seek. Use -sseof with -update 1, then join Veo 3 and Sora 2 clips with no seam.

TikTok's AI generated label isn't automatic. Set it via API
TikTok's AI generated label isn't automatic: set it with is_aigc in the Direct Post call, then burn a visible disclosure that survives your re-encode.
Ready to process videos at scale?
Start using FFmpeg Micro's simple API today. No infrastructure required.
Get Started Free