Transcription API to SRT isn't a format swap. It's timing math.

A transcription API hands you JSON: words and segments with start and end times in floating-point seconds. Your video needs a subtitle file with zero-padded timestamps, readable line breaks, and burned-in pixels. The gap is forty lines of code most pipelines get subtly wrong, and the symptom is captions a second late or an SRT FFmpeg silently refuses to render.
Quick answer: Going from a transcription API to SRT takes three moves: convert each segment's start and end seconds intoHH:MM:SS,mmmwith a comma before the milliseconds, group words into cues of at most two 42-character lines, and separate every cue with a blank line. Then burn them on withffmpeg -i input.mp4 -vf "subtitles=captions.srt" output.mp4. If you'd rather not host FFmpeg or hand-roll the timing math, FFmpeg Micro's/v1/transcribeendpoint returns a ready SRT and a transcode job burns it onto the video, both single API calls.
Microsoft made transcription cheap. The burn-in step is now the bottleneck.
Conventional wisdom says accurate captions need a good speech model. Most pipelines work that way and the output is still wrong. The failure isn't recognition quality, it's the serialization nobody tests: an ASR model returns timings relative to the audio you sent, and a video player needs cues relative to the video you're burning.
Microsoft AI shipped MAI-Transcribe-2-Streaming on October 1, 2026: 60 languages with automatic detection, a reported 2.5% word error rate, a number-one debut on Artificial Analysis, and access through the Realtime API, the Azure Speech SDK, Azure Voice Live, and Vercel. It costs $0.54 per audio hour through the end of 2026, which Microsoft also quotes as $9.00 per 1,000 minutes, about nine cents for a ten-minute video.
At nine cents, transcription isn't the hard part of a caption pipeline. The four steps around it are.
Step 1: extract audio in the format the model wants
Speech models want mono audio at a modest sample rate, not a 1080p MP4. The full video wastes upload time and can hit a request size cap. Strip to 16 kHz mono PCM and a ten-minute video becomes about 19 MB instead of 200.
The raw command:
ffmpeg -i input.mp4 -vn -ac 1 -ar 16000 -c:a pcm_s16le audio.wav
The same job as one API call, no local binary:
{
"inputs": [{ "url": "https://cdn.example.com/input.mp4" }],
"outputFormat": "wav",
"options": [
{ "option": "-vn" },
{ "option": "-c:a", "argument": "pcm_s16le" },
{ "option": "-ar", "argument": "16000" },
{ "option": "-ac", "argument": "1" }
]
}
POST that to https://api.ffmpeg-micro.com/v1/transcodes with your API key and a job id comes back immediately. More on the 16 kHz mono target: 16kHz Mono WAV for Whisper, Without a Local FFmpeg Binary. If you trimmed with -ss, note the offset now, because it bites in step 3: the transcript starts at zero no matter where the audio came from.
Step 2: keep the raw word timings, not the pretty transcript
Transcription endpoints usually return two things, and the one that looks useful is wrong. A joined text field is easy to log and useless for captions. You need the array of words or segments with start and end in seconds; concatenating it into a paragraph destroys those numbers.
Store the array as-is. If you chunked a long recording to stay under a size limit, store each chunk's offset with it, because chunk timings restart at zero: chunk three's "0.4 seconds" is really 240.4 seconds in. Streaming models are worse: partials get revised when the final result lands, so take the final timings only.
Inside a workflow tool, the size limit breaks first. Large video transcription in n8n fails. Send a URL, not the file covers that wall.
Step 3: turn word timings into a valid SRT
Converting word timings into cues is the step tutorials skip and production punishes. An SRT cue is four parts in fixed order: a sequential integer, a timestamp line, one or two lines of text, a blank line. The timestamp format is HH:MM:SS,mmm --> HH:MM:SS,mmm: two-digit hours, a comma before the milliseconds, three millisecond digits, one space each side of the arrow. Write 0:00:01.5 instead of 00:00:01,500 and parsers skip the cue or drop the file.
Grouping matters as much as formatting. Netflix's English timed-text style guide caps subtitle lines at 42 characters, a good default on a 1080p frame. Two lines maximum, one cue per breath, minimum duration near a second so short words don't flash.
The converter, overlap nudge and two-line wrap included:
const MAX_CHARS = 42; // per line
const MAX_GAP = 0.6; // silence that ends a cue
const MIN_DUR = 1.0;
const MAX_DUR = 7.0;
const pad = (n, w = 2) => String(n).padStart(w, "0");
function toSrtTime(s) {
const ms = Math.round(s * 1000);
return `${pad(Math.floor(ms / 3600000))}:${pad(Math.floor(ms / 60000) % 60)}:` +
`${pad(Math.floor(ms / 1000) % 60)},${pad(ms % 1000, 3)}`;
}
function wrap(text) {
if (text.length <= MAX_CHARS) return text;
const mid = Math.floor(text.length / 2);
let i = text.lastIndexOf(" ", mid);
if (i < 1) i = text.indexOf(" ", mid);
return i < 1 ? text : text.slice(0, i) + "\n" + text.slice(i + 1);
}
function buildCues(words, offset = 0) {
const cues = [];
let cur = null;
for (const w of words) {
const start = w.start + offset;
const end = w.end + offset;
const full = cur && cur.text.length + 1 + w.word.length > MAX_CHARS * 2;
const gap = cur && start - cur.end > MAX_GAP;
const over = cur && end - cur.start > MAX_DUR;
if (!cur || full || gap || over) {
cur = { start, end, text: w.word };
cues.push(cur);
} else {
cur.end = end;
cur.text += " " + w.word;
}
if (/[.!?]["')]?$/.test(w.word)) cur = null;
}
return cues;
}
function toSrt(cues) {
let prevEnd = 0;
return cues.map((c, i) => {
const start = Math.max(c.start, prevEnd + 0.001);
const end = Math.max(c.end, start + MIN_DUR);
prevEnd = end;
const text = wrap(c.text.trim()).replace(/\r/g, "").replace(/\n{2,}/g, "\n");
return `${i + 1}\n${toSrtTime(start)} --> ${toSrtTime(end)}\n${text}\n`;
}).join("\n");
}
Pass offset the chunk start in seconds, write UTF-8 without a BOM, end the file with a newline. The prevEnd + 0.001 line does real work: overlapping cues make libass stack two captions on top of each other, and streaming word timings overlap often.
ASS is the other option, a different format, not a prettier SRT. ASS timestamps are H:MM:SS.cc, single-digit hour and centiseconds, so you truncate milliseconds instead of reusing SRT strings. Line breaks are \N in the dialogue text, not real newlines. Curly braces are override blocks, so a transcript containing {laughs} vanishes from the frame unless you escape them. Use ASS for per-word \k karaoke highlighting or precise positioning, SRT for everything else.
Step 4: burn the captions onto the video
Burning in is one filter, and the only hard part is escaping the filename. The subtitles filter needs libass, which mainstream distro packages and the johnvansickle static builds include. Run ffmpeg -filters | grep subtitles to check yours; if it's missing, rebuild with --enable-libass or send the job to an API that has it.
ffmpeg -i input.mp4 \
-vf "subtitles=captions.srt:force_style='Fontname=Arial,Fontsize=22,PrimaryColour=&H00FFFFFF,OutlineColour=&H80000000,BorderStyle=3,Alignment=2,MarginV=60'" \
-c:v libx264 -crf 20 -preset medium -c:a copy output.mp4
Colors in force_style are ASS format, &HAABBGGRR, so &H00FFFFFF is opaque white and &H80000000 half-transparent black. BorderStyle=3 draws an opaque box behind the text, which keeps captions legible over bright footage. Alignment=2 is bottom center. force_style only overrides the Default style, so an ASS file whose lines use their own named styles will look like it ignored you.
Same filter, one request body. -filter_complex isn't supported, so use -vf:
{
"inputs": [{ "url": "gs://<YOUR_BUCKET>/1234567890-video.mp4" }],
"outputFormat": "mp4",
"options": [
{ "option": "-vf", "argument": "subtitles='https://storage.googleapis.com/.../captions.srt':force_style='Fontsize=22,BorderStyle=3,Alignment=2,MarginV=60'" },
{ "option": "-c:v", "argument": "libx264" },
{ "option": "-crf", "argument": "20" }
]
}
Poll GET /v1/transcodes/:id until status is completed, then GET /v1/transcodes/:id/download for a signed URL. A ten-minute input bills as 60 tokens, metered per second, no rounding up to the minute, and a failed job is never billed. The free tier's 200 tokens cover about 33 minutes, three ten-minute videos before you decide.
Four steps, or one call
Wiring MAI-Transcribe-2-Streaming yourself is right when you need its language coverage or streaming partials: $0.54 per audio hour plus your own extraction, converter, and libass build. It's wrong when you just want captions on a file.
To skip steps 1 through 3, /v1/transcribe takes a media URL and produces the SRT, and /v1/transcribe/:id/download returns a signed HTTPS URL you can drop into the subtitles='…' filter of the next transcode job. That URL has a ten-minute TTL, so chain the two jobs rather than storing the link. The automated version of the whole chain is the Auto Captions blueprint: upload a video, review the transcript, get the captioned file back.
Common pitfalls
Captions consistently early or late almost always means a forgotten offset. Audio extracted with -ss 00:00:30 gives a transcript starting at zero, and burning that onto the untrimmed video shifts every cue thirty seconds early. Trim both streams identically or add the offset in your converter.
Drift that grows across the video is a different bug, a sample-rate mismatch between what you declared and what you sent. If timings are right at the start and 400 ms off by minute ten, check the WAV header against the -ar you asked for.
Escaping breaks the filter more often than the file. A literal colon ends the option in a filter string, so Windows paths need subtitles=C\\:/Users/me/captions.srt, and a filename with a quote or comma needs the value wrapped and escaped. The workaround that always holds: keep the SRT beside the video with a plain ASCII name.
The subtitles filter also re-encodes the video, because it's painting pixels. -c:v copy with -vf is a contradiction FFmpeg rejects. Keep -c:a copy so audio isn't re-encoded, and expect output a generation older than the source.
When burning in is the wrong call
Burned-in captions are permanent, which makes them wrong for anything a viewer should be able to turn off or switch language on. For a player you control, ship the SRT or WebVTT as a sidecar track. For YouTube, upload the SRT as a caption track and keep the clean master. Burn in for platforms that ignore caption tracks in-feed: TikTok, Reels, and Shorts.
Captions that animate per word with template-driven motion design are a job for a template-editor service, not a filter graph. And if your pipeline regenerates narration with a TTS model, duration drift breaks your timings before the SRT does. Replace audio in video by API: the mux is the easy part covers that failure.
FAQ
Why are my captions a second early even though the transcript is correct?
Transcript timings are relative to the audio you sent, not to the video you're burning. If you trimmed, chunked, or offset the audio during extraction, add that offset back to every start and end before writing the SRT.
Can I burn captions on a video without writing an SRT file myself?
Yes. FFmpeg Micro's POST /v1/transcribe takes a media URL and returns an SRT job, and that SRT's signed download URL goes straight into the subtitles='…' filter of a transcode job, so no subtitle file lands on your disk. The Auto Captions blueprint does both steps plus a review pass.
What does it cost to caption a ten-minute video this way?
Transcription through MAI-Transcribe-2-Streaming runs about nine cents for ten minutes at the $0.54 per audio hour rate Microsoft published on October 1, 2026, good through year end. Burning the result on through FFmpeg Micro bills on input duration, 60 tokens for ten minutes, and a new free account starts with 200 tokens, roughly 33 minutes of video, no card.
To run the whole chain unattended, start with the Auto Captions blueprint and sign up free to caption your first video on the account's 200 tokens.
About Javid Jamae
Founder & CEO at FFmpeg Micro
Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.
You might also like

gpt-4o-transcribe retirement: fix captions, not the model
The gpt-4o-transcribe retirement hits October 15, 2026 with no replacement listed. Two migration paths, plus the API chain that burns captions into video.

How to use FFmpeg in n8n without touching your Docker image
Most FFmpeg n8n Docker recipes break on n8n 2.x. Compare three working paths: the runners image plus fonts, a static-ffmpeg copy, and an HTTP call on Cloud.

Large video transcription in n8n fails. Send a URL, not the file
n8n large video transcription fails because n8n holds the whole file in memory. Hand the Drive URL to an audio-extraction API and send only a 4 MB track.
Skip the command line
The Auto Captions blueprint transcribes your video and burns the captions in. You just review the transcript.
Run it (free)