ffmpegaudiotranscription

16kHz Mono WAV for Whisper, Without a Local FFmpeg Binary

·Javid Jamae·10 min read
16kHz Mono WAV for Whisper, Without a Local FFmpeg Binary

Every local Whisper build wants the same thing: 16 kHz, mono, 16-bit PCM. Every tutorial hands you the one-liner and stops there, which leaves you to figure out the part that actually breaks. That's getting the track out of an MP4 or an m4a, picking the right channel when there are two mics, and running any of it inside a runtime that has no FFmpeg binary.

Quick answer: The FFmpeg command for a 16kHz mono WAV is ffmpeg -i input.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 output.wav, which strips video, resamples to 16,000 Hz, downmixes to one channel, and writes signed 16-bit PCM. That is the exact format whisper.cpp, faster-whisper, and the openai/whisper Python package all convert to internally, so feeding anything higher buys no accuracy. If your runtime has no FFmpeg binary, like n8n Cloud, a Vercel function, or a Supabase Edge Function, send the same arguments to FFmpeg Micro as one API call and get the WAV back instead of hosting an encoder.

16 kHz is the ceiling, not a compromise

Whisper's front end resamples whatever you give it down to 16 kHz before it computes a log-mel spectrogram, so a 48 kHz master and a 16 kHz WAV reach the model as the same array of numbers. The openai/whisper GitHub discussion titled "Optimal sample rate for input audio?" (#870) settles this: there is no accuracy gain above 16 kHz, because the model never sees anything above it.

You can read the behavior straight out of the reference implementation. whisper/audio.py shells out to FFmpeg with -f s16le -ac 1 -acodec pcm_s16le -ar 16000 and reads the raw bytes off stdout. The library you're trying to feed is already running your command for you, just slower and once per file.

What this means in practice: resampling up to 44.1 kHz because "higher quality" is wasted bytes, and downsampling to 16 kHz is lossless as far as the model is concerned. Nyquist puts the ceiling at 8 kHz of audio bandwidth, and speech intelligibility lives almost entirely under that. FFmpeg's resampler applies its own anti-alias filter, so you don't need to add a manual lowpass.

The one build that genuinely requires the format is whisper.cpp. Its command-line example reads WAV through dr_wav and rejects anything that isn't 16 kHz, so the conversion isn't an optimization there, it's the entry fee. faster-whisper is more forgiving because it decodes through PyAV, but handing it a 16 kHz mono WAV skips a decode pass on every run.

Pulling the audio out of a video file

Extraction and resampling are the same command, not two steps. The flags that matter on a video input are -vn to drop the picture and -map when the file carries more than one audio track.

# MP4 with a single audio track
ffmpeg -i interview.mp4 -vn -acodec pcm_s16le -ar 16000 -ac 1 interview.wav

# m4a / AAC straight from a voice recorder
ffmpeg -i memo.m4a -acodec pcm_s16le -ar 16000 -ac 1 memo.wav

# MKV where track 0 is 5.1 commentary and track 1 is the dialogue mix
ffmpeg -i film.mkv -vn -map 0:a:1 -acodec pcm_s16le -ar 16000 -ac 1 film.wav

-acodec and -c:a are the same flag; the older spelling is what you'll see in most Whisper tutorials, so it's worth recognizing both. -vn is belt and braces for a .wav output since the RIFF container won't take H.264 anyway, but people copy these lines into jobs that write MP4s, and there it earns its place.

Then check the result before you spend GPU time on it:

ffprobe -v error -show_entries stream=codec_name,sample_rate,channels \
  -of default=noprint_wrappers=1 interview.wav

You want codec_name=pcm_s16le, sample_rate=16000, channels=1. Anything else and your transcriber will either refuse the file or quietly resample it again.

The size math, and where WAV is the wrong answer

A 16 kHz mono 16-bit WAV runs exactly 32,000 bytes per second of audio: 16,000 samples times 2 bytes times 1 channel. That number is worth memorizing, because it makes every upload-limit question arithmetic instead of guesswork.

FormatPer secondPer minutePer hour
16 kHz mono PCM s1632 kB1.92 MB115 MB
44.1 kHz stereo PCM s16176 kB10.6 MB635 MB
48 kHz stereo PCM s16192 kB11.5 MB691 MB
16 kHz mono Opus @ 24 kbps3 kB0.18 MB10.8 MB

Converting a 48 kHz stereo master to 16 kHz mono cuts it by six times, which is real. It is also still uncompressed, and that's where people get burned: OpenAI's hosted transcription endpoint caps requests at 25 MB, and 25 MB of 16 kHz mono WAV is about 13 minutes of audio. Send a one-hour interview as WAV and you'll get a 413 for a file that would have fit twice over as Opus. WAV at 16 kHz mono is the format for local Whisper builds. For the hosted API, compress instead, and if you're hitting that wall already, the Whisper 25MB limit isn't a time limit covers the re-encode side of it.

Downmix or pick a channel

-ac 1 doesn't select a channel, it averages them, and on a two-mic interview that's usually what you want: both speakers land in one transcript with no extra passes. It stops being what you want in two specific cases.

The first is phase. If one mic was wired or flipped against the other, (L + R) / 2 partially cancels the overlap and you get a thin, hollow track that transcribes badly for reasons that look like model failure. Listen to the mono downmix once before you blame Whisper.

The second is when each speaker sits on their own channel and you want them separated. Split them instead of averaging:

# One channel only
ffmpeg -i interview.wav -af "pan=mono|c0=c0" -ar 16000 -acodec pcm_s16le left.wav

# Both channels, separately, in one pass
ffmpeg -i interview.wav -filter_complex "[0:a]channelsplit=channel_layout=stereo[l][r]" \
  -map "[l]" -ar 16000 -acodec pcm_s16le host.wav \
  -map "[r]" -ar 16000 -acodec pcm_s16le guest.wav

For 5.1 sources, dialogue lives in the center channel and the default downmix matrix mixes it in alongside music and effects. -af "pan=mono|c0=FC" takes center alone and often produces a noticeably cleaner transcript than -ac 1 on the same film.

When the runtime has no FFmpeg binary

The conversion is trivial on a laptop and awkward everywhere a pipeline actually runs. n8n disabled the Execute Command node by default in v2.0 for security reasons, so the old "just shell out" workaround is gone, and the official n8n Docker image went distroless in 2026, which breaks every apk add ffmpeg recipe you'll find in the forums. Vercel functions and Supabase Edge Functions have no binary and no subprocess to call one with. Three different runtimes, same dead end.

The fix isn't a smaller static FFmpeg build squeezed into your bundle. It's not running FFmpeg in that runtime at all. FFmpeg Micro takes the same arguments over HTTP: you post an input URL and the FFmpeg arguments, get a job ID back, then poll or take a webhook and download the WAV.

-vn -acodec pcm_s16le -ar 16000 -ac 1

That's the whole difference between the two paths. Same flags, no encoder to host, no binary to keep current, and it works the same from an n8n HTTP Request node, a Make scenario, a Vercel function, or an MCP tool call inside an agent loop. The request and response shapes are in the docs. If your source audio already lives in object storage, pass the URL rather than the bytes. Give FFmpeg a presigned URL explains why routing a 600 MB master through your workflow engine is the thing that actually falls over.

Pitfalls that cost people an afternoon

Piping WAV to stdout produces a broken header. FFmpeg backfills the RIFF size field after it finishes writing, and it can't seek on a pipe, so the header ends up with a placeholder length. Some readers cope, dr_wav in whisper.cpp does not. Write to a real file, or use -f s16le raw PCM if your consumer accepts it.

Stream selection picks the audio track with the most channels, not the one you meant. A file with a 5.1 commentary track and a 2.0 dialogue mix hands you the commentary. Be explicit with -map 0:a:0.

A video-only MP4 gives you Output file #0 does not contain any stream rather than an empty WAV. The neighboring failure, output file is empty, nothing was encoded, shows up when a -map points at a stream that isn't there.

Converting an already-16 kHz file is a no-op worth skipping. Check with ffprobe first; on a batch of thousands, skipping the conversion on files that are already compliant is the cheapest speedup available.

Finally, don't normalize or denoise "to help accuracy" without measuring it. Whisper was trained on messy audio and aggressive noise reduction removes cues the model uses. Get the format right first, then test whether anything else earns its place.

FAQ

Does a higher sample rate make Whisper more accurate?

Higher sample rates make no difference to Whisper's accuracy, because the model resamples everything to 16 kHz before computing its mel spectrogram. A 48 kHz WAV and a 16 kHz WAV of the same recording produce identical input to the model and identical transcripts, with the 48 kHz file just taking longer to load.

Do I need a WAV file, or will an MP3 work?

Whether you need WAV depends on which build you're running. whisper.cpp's command-line example reads WAV only and requires 16 kHz, so MP3 won't load at all. The openai/whisper Python package and faster-whisper both decode MP3, m4a, and MP4 fine, but they call FFmpeg internally to do it, so pre-converting to 16 kHz mono WAV just moves that work out of your inference loop.

How do I convert an MP4 to 16 kHz mono WAV without installing FFmpeg?

You send the conversion to an API instead of running a binary locally. FFmpeg Micro takes -vn -acodec pcm_s16le -ar 16000 -ac 1 as job arguments over HTTP and returns the WAV, which is the only route available in runtimes like n8n Cloud, Vercel functions, and Supabase Edge Functions where there's no subprocess to shell out to.

Should I use -ac 1 or pick a single channel for a two-mic interview?

Use -ac 1 for a two-mic interview unless you need the speakers kept apart. The downmix averages both mics into one track and transcribes as a single conversation, which is usually the goal. Use channelsplit to write two WAVs when you want per-speaker transcripts, or pan=mono|c0=c0 when one channel is unusable.

Why does whisper.cpp reject my WAV file?

whisper.cpp rejects WAV files that aren't 16 kHz, and also chokes on WAVs written to a pipe, because the RIFF header carries a placeholder size that FFmpeg never got to backfill. Run ffprobe on the file and confirm pcm_s16le, 16000, and 1 channel before assuming the model is at fault.

Extract and resample in one call, from code, from n8n, or from an agent, with no encoder to host and a free tier to test it on: sign up free and run your first conversion against a real interview file.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Ready to process videos at scale?

Start using FFmpeg Micro's simple API today. No infrastructure required.

Get Started Free