subtitlesn8n

gpt-4o-transcribe retirement: fix captions, not the model

··10 min read
gpt-4o-transcribe retirement: fix captions, not the model

On October 15, 2026, Azure Foundry retires gpt-4o-transcribe version 2025-03-20. If that model is the first step in your caption workflow, every call after that date comes back 410 Gone, and the Replacement column on Microsoft's own schedule is empty.

Quick answer: The gpt-4o-transcribe retirement takes effect October 15, 2026, when Azure Foundry removes gpt-4o-transcribe (2025-03-20), gpt-4o-mini-transcribe (2025-03-20) and gpt-4o-mini-tts (2025-03-20) from service and all inference requests return 410 Gone. You have two fixes: redeploy on a dated successor that is still supported, such as gpt-4o-mini-transcribe 2025-12-15 which runs to June 15, 2027, and keep maintaining your own SRT and burn-in steps; or chain POST /v1/transcribe and POST /v1/transcodes on FFmpeg Micro, which transcribes the audio and burns the subtitles into the frame so the thing you get back is a captioned MP4 instead of a transcript.

What retires on October 15, and what the empty Replacement column means

Microsoft's model retirement schedule was last updated September 23, 2026, and four rows on it matter to anyone running captions or AI clip generation through Azure.

ModelVersionLifecycleRetirement dateReplacement
gpt-4o-transcribe2025-03-20GA2026-10-15None listed
gpt-4o-mini-transcribe2025-03-20GA2026-10-15None listed
gpt-4o-mini-tts2025-03-20Preview2026-10-15None listed
sora-22025-12-08Preview2026-10-15None listed

That blank last column is the part people misread. Azure does auto-upgrade Global Standard, Data Zone Standard and Standard deployments when a version retires, but it upgrades them to the declared replacement. No replacement is declared for gpt-4o-transcribe 2025-03-20, so there is nothing for the service to move your deployment onto. The deployment stops serving. Microsoft's lifecycle policy is explicit about the end state: retired models can't accept new deployments, existing deployments stop working, and "all inference requests return 410 Gone." The support FAQ is equally blunt about asking for more time: "Retirement dates aren't extendable."

Worth checking your own deployment config while you're in there. If versionUpgradeOption is set to NoAutoUpgrade, the deployment stops working at retirement even for models that do have a successor.

And if your plan was to fall back to Azure's hosted Whisper, read the date first: whisper version 001 is on the same page with a retirement date of December 15, 2026. That buys you ten weeks, not a year.

Path one: redeploy on a dated successor, and re-test accuracy

Moving to a newer dated model is the fastest change to make this week, but it isn't the clean version bump the phrase suggests, because the still-supported transcription models aren't the same model you were calling. The two realistic targets on Azure are gpt-4o-mini-transcribe 2025-12-15, supported to June 15, 2027, and gpt-4o-transcribe-diarize 2025-10-15, supported to April 15, 2027. The first is a smaller variant. The second adds speaker labels and a different response shape.

Either way you're changing deployment names, re-running your worst audio through it, and checking that your timestamp parsing still holds. That is an afternoon, not a minute. Microsoft's own migration guidance says the same thing in a politer register: evaluate candidates against your own prompts and representative data rather than public benchmarks.

The other thing path one doesn't fix is everything downstream. A transcript is not a caption. The reason the retirement hurts is that a vendor swap at step one ripples through steps two, three and four, and those steps are yours to maintain.

Path two: make the whole caption job one chain you don't maintain

Captions are three jobs stacked, and only the middle one just got retired. Pull the audio out of the video, turn the transcript into timed cues, then render those cues into pixels. Swapping the transcription vendor leaves the other two exactly where they were, which is the argument for putting all three behind one API instead of re-plumbing the first one every time a model hits an 18-month lifecycle wall.

Step one: get audio out of the video

FFmpeg does this in one command, and it cuts upload size before transcription:

ffmpeg -i input.mp4 -vn -c:a libmp3lame -q:a 2 audio.mp3

The same operation as an API call, no binary required:

curl -X POST https://api.ffmpeg-micro.com/v1/transcodes \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": [{ "url": "https://cdn.example.com/input.mp4" }],
    "outputFormat": "mp3",
    "options": [
      { "option": "-vn" },
      { "option": "-c:a", "argument": "mp3" },
      { "option": "-b:a", "argument": "192k" }
    ]
  }'

If your transcriber is a local Whisper build rather than a hosted endpoint, it wants 16 kHz mono PCM, which is a different flag set. We covered that encode in 16kHz mono WAV for Whisper, without a local FFmpeg binary.

Step two: transcribe to timed SRT

POST /v1/transcribe takes a Cloud Storage path and returns a job you poll:

{
  "media_url": "gs://your-bucket/1234567890-audio.mp3",
  "language": "en",
  "task": "transcribe"
}

Omit language and it detects automatically. Set task to translate instead of transcribe and you get English cues from non-English audio. Poll GET /v1/transcribe/:id until status is completed, then GET /v1/transcribe/:id/download returns a signed HTTPS URL to the SRT, valid for ten minutes.

Step three: burn the cues into the frame

Line-level subtitles rendered into the video are what social platforms actually show. The raw command:

ffmpeg -i input.mp4 \
  -vf "subtitles=captions.srt:force_style='Alignment=2,Fontsize=22'" \
  -c:a copy output.mp4

The one-call version passes the signed SRT URL straight into the same filter through POST /v1/transcodes, so the SRT never touches your disk:

{
  "inputs": [{ "url": "gs://your-bucket/input.mp4" }],
  "outputFormat": "mp4",
  "options": [
    { "option": "-vf", "argument": "subtitles='SIGNED_SRT_URL':force_style='Alignment=2,Fontsize=22'" },
    { "option": "-c:a", "argument": "copy" }
  ]
}

Alignment=2 is bottom center, 8 is top center, 1 and 3 are corners. Raise Fontsize for vertical video, where 22 reads small at 1080x1920.

The n8n and Make version of the same chain

In n8n or Make this is six HTTP nodes and no container changes, which matters because n8n Cloud gives you no way to install an FFmpeg binary at all. The order:

  1. Upload the video with POST /v1/upload/presigned-url, PUT the bytes directly to storage, then POST /v1/upload/confirm. The confirm response hands back the duration.
  2. Submit the audio extraction, or skip it and transcribe the video file directly.
  3. POST /v1/transcribe with the gs:// URL.
  4. Poll GET /v1/transcribe/:id on a short loop until completed.
  5. GET /v1/transcribe/:id/download for the signed SRT URL.
  6. POST /v1/transcodes with the subtitles= filter, poll it, then GET /v1/transcodes/:id/download.

Three calls plus the uploads, and the output of the last one is a finished captioned MP4. The full n8n build, including the sub-workflows for upload and transcode polling, is walked through in How to use FFmpeg with n8n Cloud, and the API reference for every endpoint above is in the docs. Billing is on the duration of media you send, metered per second with no rounding up to the minute, and failed jobs are never billed. A new account gets 200 tokens, about 33 minutes of video, with no card.

Common pitfalls

The signed SRT URL expires in ten minutes, so a Wait node or a human approval step between step five and step six will hand FFmpeg a dead URL. Fetch the SRT URL immediately before you submit the transcode, not at the top of the workflow.

A few more that cost real debugging time:

  • 410 Gone is permanent, and retry logic treats it like a blip. An n8n node with three retries and exponential backoff will burn through its attempts and then fail anyway, so check for the status explicitly and alert instead of retrying.
  • Music-only audio transcribes to an empty SRT, and a heavy music bed under narration produces cues that are confidently wrong. Gate on transcript length before you render.
  • A signed URL's query string contains :, & and =, all of which FFmpeg's filter parser treats as delimiters. Quote the URL inside the filter argument and escape colons if you build the filter string yourself.
  • -c:a copy keeps the audio untouched, which is what you want. Dropping it re-encodes the audio for no reason and adds a generation of loss.

One pitfall is calendar-shaped rather than technical. If the same workflow generates clips through Azure's sora-2, that endpoint dies on October 15 too, and its Replacement column is empty for a harder reason: there is no successor anywhere. Preview models with no replacement get 30 days notice and then return 410 Gone. We wrote up what to do with those steps in A model swap isn't the Sora API shutdown alternative for n8n.

When this isn't the right approach

Burned-in captions are permanent, and that's the trade. If your editors need to read the transcript and fix names before anything renders, don't automate the render: pull the raw SRT from GET /v1/transcribe/:id/download, edit it wherever you edit text, and feed the corrected file to the transcode. There's no caption editor UI here, because FFmpeg Micro is an API rather than an editor.

Two other cases call for something else. Word-by-word karaoke highlighting is a different animation problem than line-level cues, and a template-editor service built around keyframed text will get you there faster. Live captioning on a stream isn't a batch job at all, and nothing described above applies to it.

FAQ

What happens if I call gpt-4o-transcribe after October 15, 2026?

Requests to a retired Azure Foundry model return 410 Gone. There's no grace period, no silent fallback to a newer version, and retirement dates are not extendable on request.

Is there a drop-in replacement for gpt-4o-transcribe on Azure?

Azure lists no replacement for gpt-4o-transcribe 2025-03-20. The nearest supported options are gpt-4o-mini-transcribe 2025-12-15, which runs to June 15, 2027, and gpt-4o-transcribe-diarize 2025-10-15, which runs to April 15, 2027 and adds speaker labels, so both need accuracy testing against your own audio before you cut over.

Can I keep using Azure's hosted Whisper instead?

Azure's whisper version 001 has a retirement date of December 15, 2026, two months after the gpt-4o transcribe models. It works as a stopgap and not as a migration target.

How do I burn an SRT into a video without installing FFmpeg?

Send the video and the SRT URL to a hosted FFmpeg API and let it run the subtitles= filter server-side. On FFmpeg Micro that's POST /v1/transcodes with -vf subtitles='URL':force_style='Alignment=2,Fontsize=22', and the output is an MP4 with the captions rendered into the pixels.

Does this work in Make and Zapier too?

The chain is plain HTTP calls, so Make, Zapier and n8n all run it with their generic HTTP modules, and the same requests work from Python, Node or an AI agent through the MCP server.

If you'd rather test the output before you rewire anything, the Auto Captions blueprint is this exact chain behind one upload: pick a language, pick a caption style, download a captioned MP4. Start on the free tier and run one of your real clips through it before the 15th, while you still have a working transcription endpoint to compare against.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Skip the command line

The Auto Captions blueprint transcribes your video and burns the captions in. You just review the transcript.

Run it (free)