A model swap isn't the Sora API shutdown alternative for n8n

Your n8n workflow broke yesterday. OpenAI removed the Videos API and every sora-2 model ID on September 24, 2026, and if your scenario lived in Make, its OpenAI video modules vanished six days earlier. Every migration guide you'll find tells you which model to use instead and then stops at the moment the new clip URL appears, which is exactly where the real work starts.
Quick answer: The working Sora API shutdown alternative is an async generate-and-poll branch: replace the Sora node with an HTTP Request POST to Veo 3.1 (predictLongRunningon the Gemini API) or Kling 3.0, then a Wait node plus a status GET that loops untildoneis true. That gets you a file, but not a Sora-shaped one. Veo 3.1 renders only 16:9 or 9:16 in 4, 6, or 8 second clips with audio already baked in, so anything downstream that assumed a fixed length, ratio, or silent track now needs an explicit normalization pass. You can run that pass in FFmpeg yourself, or send the clip to FFmpeg Micro as one API call per step: reframe to 1080x1920, concat the clips, burn captions, mix the music bed. No encoder to host, no timeouts to babysit, free tier to start.
What actually shut down, and when
OpenAI's deprecations page lists the Videos API and all five Sora identifiers (sora-2, sora-2-pro, sora-2-2025-10-06, sora-2-2025-12-08, sora-2-pro-2025-10-06) as removed on 2026-09-24, announced on 2026-03-24. The replacement column is blank. There's no successor model to point a workflow at, which is why this migration is a pipeline change rather than a string swap.
Make retired its OpenAI video modules on 2026-09-18, ahead of the API itself, and told users to update scenarios before the 24th. If your scenario still shows those modules, they're decoration.
Before you trust a template: n8n's gallery still hosts workflow 11276, "Generate & publish AI videos with Sora 2, Veo 3.1, Gemini & Blotato," whose Sora 2 Pro branch calls dead API surface. Popular does not mean current.
Swap the Sora node for generate-and-poll
Sora's Videos API behaved close enough to synchronous that most workflows treated it as one node in, one file out. Veo and Kling return an operation handle instead and make you poll it, so the single node becomes a three-node branch: POST, Wait, GET, looping back to Wait until the job reports done.
For Veo 3.1 through the Gemini API, the HTTP Request node posts to the long-running endpoint:
{
"method": "POST",
"url": "https://generativelanguage.googleapis.com/v1beta/models/veo-3.1-generate-preview:predictLongRunning",
"headers": { "x-goog-api-key": "{{ $credentials.geminiApiKey }}" },
"body": {
"instances": [{ "prompt": "{{ $json.prompt }}" }],
"parameters": {
"aspectRatio": "9:16",
"durationSeconds": 8,
"resolution": "1080p"
}
}
}
The response carries an operation name. Feed it to a Wait node set to 10 seconds (the interval Google's own docs recommend), then a GET on https://generativelanguage.googleapis.com/v1beta/{{ $json.name }}, then an IF node on done that loops back to Wait when false. The finished sample URL lands at response.generateVideoResponse.generatedSamples[0].video.uri.
Two Veo constraints will bite you if you copy a 720p template: 1080p and 4K output require the 8-second duration, and the only aspect ratios are 16:9 and 9:16. Kling 3.0 is the looser option at 3 to 15 seconds (5s default) and adds 1:1, with native audio on by default.
Add a download step right after the poll succeeds. Google stores generated videos on its servers for two days and then removes them, so a workflow that hands the Google URL straight to a publisher and retries tomorrow will 404. Pull the file into S3, Google Drive, or your own bucket inside the same execution.
If a polling loop feels wasteful on long generations, the same branch works webhook-first. The pipeline patterns for long-running video jobs apply unchanged here.
The part the migration guides skip
A Sora replacement gives you a clip, not the clip your pipeline expected. Your old workflow was calibrated to one generator's output shape, and every assumption baked into the nodes after it (fixed duration, a known resolution, a silent or replaceable audio track) is now wrong. The four steps below put the output back into a predictable shape, and each is a single job you can run against FFmpeg Micro's API from an HTTP Request node instead of shelling out to a binary n8n's distroless image no longer contains.
One envelope covers all four. Here's the reframe as a full request; the other three are the same POST with a different command.
{
"method": "POST",
"url": "https://api.ffmpeg-micro.com/v1/jobs",
"headers": { "Authorization": "Bearer {{ $credentials.ffmpegMicroKey }}" },
"body": {
"input": "{{ $json.videoUrl }}",
"command": "-vf scale=1080:1920:force_original_aspect_ratio=decrease,pad=1080:1920:(ow-iw)/2:(oh-ih)/2:color=black,setsar=1 -c:v libx264 -crf 20 -preset veryfast -c:a copy",
"output": "reframed.mp4"
}
}
Field names come from the docs; the job returns an ID you poll or receive by webhook, same as the generator.
Reframe to a clean 1080x1920
The reframe decides whether your clip gets letterboxed or cropped, and the two filter chains differ by one word. The command above pads: it scales the longest side to fit and fills the rest with black, which never loses pixels. To fill the frame instead, swap decrease for increase and replace pad with crop=1080:1920, which loses the edges but leaves no bars.
Use crop for 16:9 Veo output going to Reels or TikTok, where black bars read as a reposted video. Use pad when the generated composition has content near the frame edges you can't afford to cut.
Meta's 1:1 and 4:5 placements need this step no matter which model you pick, since neither Veo nor Sora ever exposed those ratios as a parameter. The same math for those two shapes is in Veo and Sora lock AI video aspect ratio. Where the reframe is the whole job, the resize-format blueprint is the one-click version: drop the file, pick 9:16, pad or crop, download.
Concatenate the clips
Multi-clip generations almost never concat with -c copy, because Veo at 8 seconds and Kling at 5 rarely agree on resolution or frame rate, and the demuxer will either refuse or produce a file that stutters at the seam. Normalize each input to one canvas first, then concat in a single filter_complex:
[0:v]scale=1080:1920:force_original_aspect_ratio=decrease,pad=1080:1920:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v0];
[1:v]scale=1080:1920:force_original_aspect_ratio=decrease,pad=1080:1920:(ow-iw)/2:(oh-ih)/2,setsar=1,fps=30[v1];
[0:a]aresample=48000[a0];[1:a]aresample=48000[a1];
[v0][a0][v1][a1]concat=n=2:v=1:a=1[v][a]
The setsar=1 and aresample=48000 look redundant until one generation comes back at 44.1 kHz and the audio drifts a frame per clip. If you want motion to continue across the cut rather than just butting clips together, seed each generation from the previous clip's final frame, a separate job covered in extract the last frame to chain Veo clips cleanly.
Burn the captions
Captions go on after the concat, never before, or the burn gets re-encoded a second time and the text softens. Point the subtitles filter at your SRT and set the style explicitly, since the defaults put text at the bottom edge where TikTok's UI covers it:
-vf "subtitles=captions.srt:force_style='FontName=Inter,FontSize=16,PrimaryColour=&H00FFFFFF,OutlineColour=&H00000000,BorderStyle=1,Outline=2,Shadow=0,Alignment=2,MarginV=180'"
MarginV=180 lifts the block above the caption and button stack on a 1920-tall frame.
Mix the music, don't overwrite it
Audio mixing is the step most migrated workflows get wrong, and it fails silently. Sora clips arrived without a usable dialogue track, so the standard pattern was -map 0:v -map 1:a, replacing audio wholesale. Veo 3.1 and Kling 3.0 both generate audio natively, so that same -map now throws away the model's dialogue and sound design and leaves you with a music-only clip that reads as broken.
Mix instead:
-filter_complex "[0:a]volume=1.0[a0];[1:a]volume=0.18[a1];[a0][a1]amix=inputs=2:duration=first:normalize=0[a]" -map 0:v -map "[a]" -c:v copy
normalize=0 matters. By default amix divides each input's gain by the number of inputs, which halves your dialogue the moment you add a bed. For music that actually gets out of the way under speech rather than sitting at a fixed 18%, sidechain ducking is the better filter.
Make scenarios: the HTTP module equivalent
Make users have no module to restore, so the fix is an HTTP "Make a request" module with the same POST body shown above, a Sleep module set to 10 seconds, and a repeater or router that re-queries the operation until done returns true. The JSON is identical; only the node names change. The FFmpeg Micro calls work the same way from Make and Zapier, since they're plain HTTP with a bearer token.
Common pitfalls
The same five mistakes come up in roughly this order:
- Hard-coding the generator. Sora got six months of notice and still left no successor. Put the model endpoint, ratio, and duration in a Set node so the next shutdown is a config edit.
- Skipping the download. Two-day retention on Veo output means a failed publish step that retries tomorrow has nothing to fetch.
- Polling with no ceiling. An IF loop with no max-iteration guard will spin on a failed operation until the execution times out. Cap it at 60 iterations and route the failure branch somewhere visible.
- Assuming a duration. Veo gives 4, 6, or 8 seconds and Kling gives 3 to 15, so any downstream timing (caption offsets, music fade points, a hard-coded
-t) has to read the actual duration rather than trust a constant. - Re-encoding at every stage. Reframe, concat, and caption each re-encode. Chain them into as few jobs as you can, and keep
-c:v copyon the audio mix since the video is untouched there.
When this isn't the right approach
An HTTP-and-FFmpeg pipeline is the right call when the output shape is predictable and the variation lives in the prompt. It's the wrong call when every render needs a different layout, timed text animations, or a designer changing the template weekly. That work belongs in a template-editor service with a visual composer, where you're paying for the layout tool rather than the encode. Same for live streaming: this is a batch job pattern, not a real-time one.
FAQ
What replaced the Sora API after the September 24 shutdown?
OpenAI named no replacement for the Sora API. The deprecations entry removed all five sora-2 identifiers with the replacement column empty, so migrating builders are choosing between Google's Veo 3.1 on the Gemini API, Kling 3.0, and Seedance 2.0 on their own merits rather than following an official path.
Can I keep my n8n workflow's structure and only change the model?
Changing the model alone won't restore an n8n workflow that used Sora. Sora's Videos API returned a file the workflow could use almost immediately, while Veo and Kling return an operation handle, so the single node becomes POST, Wait, poll, and IF. Everything downstream that assumed Sora's clip length, ratio, or silent audio track also needs revisiting.
How do I get 1:1 or 4:5 video out of Veo 3.1?
Veo 3.1 cannot output 1:1 or 4:5 at all; it exposes only 16:9 and 9:16. Generate in the closest ratio and crop or pad to the target shape afterward with a scale-and-crop pass, which is a single FFmpeg job or one API call per output ratio.
Do I still need burned-in captions if the model generates dialogue audio?
Burned-in captions still matter even when the model generates dialogue, because most social feeds autoplay muted and platform auto-captions, generated after upload, are wrong as often as not on synthetic speech. Burning your own SRT is the only way to control wording and position.
Should I run FFmpeg inside n8n instead of calling an API?
Running FFmpeg inside n8n has gotten harder, not easier. The official Docker image went distroless, so the apk add ffmpeg recipe in every older tutorial fails, and n8n v2.0 disabled the Execute Command node by default. If you want the comparison across community nodes and HTTP approaches, see which n8n FFmpeg node should you use.
Rebuilding the generate-and-poll branch is an afternoon. The normalization steps after it are four HTTP Request nodes you can paste in once and reuse against whatever model replaces Veo next year. Sign up free and run the reframe against one of yesterday's broken clips to see what comes back.
About Javid Jamae
Founder & CEO at FFmpeg Micro
Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.
You might also like

Choosing a RenderIO alternative for n8n? Look past the node
Looking for a RenderIO alternative for n8n? Compare per-command quotas, clip duration caps, and webhook gating against usage-based FFmpeg API pricing.
AI Avatar Video Automation Breaks at Assembly, Not Generation
AI avatar video automation stalls after the talking head renders. The assembly chain: stitch b-roll, burn captions, add a music bed, and ship 9:16 vertical.

The fal.ai FFmpeg API merges clips. It can't burn captions.
The fal ai ffmpeg api covers compose and merge-videos at $0.0002/second. Where that fits an AI video pipeline, and where a full FFmpeg surface wins out.
Ready to process videos at scale?
Start using FFmpeg Micro's simple API today. No infrastructure required.
Get Started Free