n8ntranscriptionworkflow-automation

Large video transcription in n8n fails. Send a URL, not the file

·Javid Jamae·10 min read
Large video transcription in n8n fails. Send a URL, not the file

A 4K field clip sitting in Google Drive at 420 MB is not a transcription problem. It becomes one the moment your n8n Cloud workflow tries to download it, hold it, and POST it to AssemblyAI in the same execution. The run dies, the log says nothing useful, and the same workflow works fine on a 30 MB screen recording.

Quick answer: n8n large video transcription fails on n8n Cloud because n8n keeps binary data in the instance's memory, so a 200-500 MB 4K clip can exhaust RAM before it ever reaches AssemblyAI, Deepgram, or Whisper. The self-hosted fix is N8N_DEFAULT_BINARY_DATA_MODE=filesystem plus your own FFmpeg binary, and n8n Cloud lets you set neither. On Cloud, skip the download: POST the video's URL to FFmpeg Micro's /v1/transcodes endpoint with -vn and a 16 kHz mono audio filter, then pass the returned audio URL to your transcription provider, so n8n only ever moves JSON.

Why n8n Cloud runs out of memory on a 200 MB clip

n8n Cloud stores binary data in the instance's memory, and there is no settings panel where you can change that. A Google Drive Download node followed by an HTTP Request node keeps at least two copies of the file alive at once: the item's binary buffer and the request body being assembled. n8n also base64-encodes binary between nodes, which inflates the payload by about a third. A 420 MB source clip is realistically 1.1 GB of heap by the time the HTTP node fires.

This is the live complaint, not a hypothetical. The n8n forum thread "How to manage large files within n8n Cloud" (community.n8n.io/t/…/287925) describes exactly this setup: 200-500 MB 4K field clips in Drive, AssemblyAI at the other end, the workflow dying in the middle. Thread 296021 from May 2026 repeats it generically, and thread 121280 has users OOMing on 1 GB files through the Read File from Disk node. The complaint has been open, in one form or another, for two years.

The old escape hatches are gone too. n8n v2.0 disables the Execute Command node by default because arbitrary shell execution is a security problem in shared environments, so you can't shell out to FFmpeg anymore. The official docker.n8n.io/n8nio/n8n image went distroless in June 2026, so the apk add --no-cache ffmpeg recipe in every tutorial now fails outright (the two fixes are here). Feature request 261143, asking n8n to just bundle FFmpeg in the image, is still open.

Pass the URL, not the file

Every service in this chain already accepts a URL, which means the video never has to enter n8n at all. Google Drive can serve a direct download link. FFmpeg Micro takes inputs[].url. AssemblyAI takes audio_url, Deepgram takes {"url": ...}. Line those up and n8n's entire job is moving three short JSON bodies between HTTP Request nodes.

That reframes the problem. You are not looking for a way to carry 420 MB through a workflow engine. You are looking for one service that can read a video URL and write a small audio URL, so nothing downstream has to carry anything. On n8n Cloud, where you can't install a binary or raise a memory ceiling, that service has to be an API call. FFmpeg Micro is built for exactly this shape, and the n8n integration page has the node-by-node wiring alongside the full request reference in the API docs.

The four-node recipe

The whole workflow is four HTTP Request nodes and a Wait node, and none of them has binary data enabled.

1. Get a direct download URL for the Drive file

Google Drive share links are HTML pages, not media. Use the Google Drive node's Share File operation to set the file to "anyone with the link," then build the direct URL as https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t. The confirm=t parameter matters: without it, Drive returns a virus-scan interstitial page for anything over roughly 100 MB, and your extraction job will download 2 KB of HTML instead of a video.

If your files live in S3 or R2 instead, a presigned URL works the same way and skips this dance entirely. The pattern is the same one in give FFmpeg a presigned URL.

2. Submit the audio extraction job

One POST creates the job. The -vn flag drops the video stream, so FFmpeg never decodes a single 4K frame, and the audio filter forces the sample rate and channel layout that speech models want.

curl -X POST https://api.ffmpeg-micro.com/v1/transcodes \
  -H "Authorization: Bearer $FFMPEG_MICRO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": [{ "url": "https://drive.usercontent.google.com/download?id=FILE_ID&export=download&confirm=t" }],
    "outputFormat": "mp3",
    "options": [
      { "option": "-vn" },
      { "option": "-af", "argument": "aformat=sample_rates=16000:channel_layouts=mono" },
      { "option": "-c:a", "argument": "libmp3lame" },
      { "option": "-b:a", "argument": "64k" }
    ]
  }'

The response comes back immediately with an id and status: "pending". In n8n, that's a plain HTTP Request node with the JSON above in the body field.

3. Poll until the job completes

Add a Wait node set to 5 seconds, then GET /v1/transcodes/{id}, then an IF node that loops back while status is pending or processing. Audio-only extraction is quick because FFmpeg skips video decoding, so most of the wall time is the API pulling your source file. Jobs with inputs longer than two minutes run in a separate long lane on their own worker, so a 40-minute interview can't get stuck behind someone else's short clip.

4. Hand the audio URL to your transcription provider

When the status flips to completed, call GET /v1/transcodes/{id}/download, which returns {"url": "https://storage.googleapis.com/..."}. That signed URL goes straight into the provider's request body:

{
  "audio_url": "{{ $json.url }}",
  "webhook_url": "https://your-instance.app.n8n.cloud/webhook/transcript-ready"
}

AssemblyAI fetches the audio itself. n8n never sees a byte of it.

The FFmpeg command this replaces

On your own machine, the equivalent is a single command, and it's worth knowing because it explains what the API options map to.

ffmpeg -i field-clip-4k.mov -vn -ac 1 -ar 16000 -c:a libmp3lame -b:a 64k audio.mp3

One difference will bite you if you translate it literally. The FFmpeg Micro option allowlist doesn't include -ar or -ac, so sending {"option": "-ar", "argument": "16000"} returns a 400 with The option -ar is not supported. Use the -af aformat=... form above instead, which reaches the same result through the audio filter graph. Filter arguments also reject ;, &&, and |, so pan=mono|c0=... won't pass either.

What the payload drops to

Audio size tracks duration, not resolution, which is why the savings are so lopsided on 4K footage. A nine-minute clip produces the same 4.3 MB MP3 whether the source was 200 MB or 500 MB.

StageWhat n8n holds in memorySize
9-minute 4K source in Drivea URL string480 MB
Extraction job requestJSON body~400 bytes
Extracted 16 kHz mono MP3a signed URL string4.3 MB
Transcription requestJSON body~200 bytes

If your provider wants lossless input, "outputFormat": "wav" gives you 16-bit PCM at 1.92 MB per minute. That's a real constraint worth doing the arithmetic on: WAV crosses OpenAI's 25 MB upload cap at about 13 minutes of audio, which is the trap behind the Whisper 25MB limit. Mono MP3 at 64 kbps doesn't hit that cap until roughly 52 minutes.

Polling or webhooks: use both, on different sides

Use polling for the extraction step and a webhook for the transcription step, because the two jobs have very different durations. FFmpeg Micro's status endpoint is what you poll, with a Wait node between checks, and the job typically settles fast enough that a 5-second interval costs you only a handful of extra executions. Transcription is the slow half: a 40-minute recording can take several minutes at any provider, and holding an n8n execution open that long for a poll loop is wasteful.

So give the provider an n8n Webhook node URL and let it call you back. The workflow splits into two: one that submits, one that receives the transcript. That's the reliable shape for anything long-running, covered in more depth in webhooks for long-running video jobs.

Common pitfalls

Most failures in this pipeline come from the URL hand-off rather than the media itself.

  • The signed download URL from /v1/transcodes/{id}/download is valid for 10 minutes. Generate it in the node immediately before the transcription call, not at the top of the workflow.
  • A Drive link that isn't shared publicly returns an HTML login page. FFmpeg Micro downloads it, then fails with an invalid-data error on a 2 KB "video."
  • URL-based inputs cap at 1 GB. Above that you'll get File too large. Maximum input file size is 1GB for URL-based inputs, and you'll need the direct upload path, which allows up to 2 GB.
  • Leaving the Google Drive Download node in the workflow "just to check the file exists" reintroduces the exact memory problem you removed. Use the Drive node's Get File metadata operation instead.
  • Setting -b:a below 32k on 16 kHz mono audio starts audibly smearing consonants, which shows up as word error rate, not as a broken file.

When this is the wrong approach

Routing around n8n makes sense when the video is large and you only need the words. Keep the file local when you actually need the frames in the same run, for example burning captions back onto the clip, because you'd be moving the video anyway. It's also unnecessary if you self-host n8n with filesystem binary mode and plenty of RAM and you process one file a week: the memory ceiling is the whole reason for the detour, and without it a self-hosted FFmpeg install is fine (community node options compared here).

And if your media already arrives as separate audio, skip extraction entirely and hand the audio URL straight to the provider. One less call is always better than a faster call.

FAQ

Can n8n Cloud handle a 500 MB video file?

n8n Cloud cannot reliably move a 500 MB video through a workflow, because binary data is held in the instance's memory and n8n base64-encodes it between nodes. Pass the file's URL between nodes instead and let an external API do the reading and writing.

Do I need to self-host n8n to transcribe large videos?

Self-hosting is not required for large video transcription. Self-hosting gives you N8N_DEFAULT_BINARY_DATA_MODE=filesystem and the ability to install FFmpeg, but a URL-in, URL-out audio extraction call removes the need for both, and it runs identically on n8n Cloud, Make, and Zapier.

What's the best way to extract audio for a transcription API?

The safest target for a transcription API is 16 kHz mono, which is what Whisper and most speech models resample to internally anyway. Use -vn to drop video, aformat=sample_rates=16000:channel_layouts=mono to set the format, and MP3 at 64 kbps to keep the file small enough for upload-based endpoints.

How long does an audio extraction job take on a 4K file?

Extraction wall time on a 4K file is dominated by downloading the source, not by encoding, since -vn means FFmpeg never decodes a video frame. Inputs longer than two minutes are routed to a dedicated long lane so they don't queue behind short jobs.

You can build the whole chain with an API key and four HTTP Request nodes, and the free tier covers enough minutes to test it against one of your real 4K clips before you wire up the rest. Sign up free and run the extraction call against a Drive URL you already have.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Ready to process videos at scale?

Start using FFmpeg Micro's simple API today. No infrastructure required.

Get Started Free