How to Add Chapters to MP4 with FFmpeg (FFMETADATA)

You have timestamps from a transcript or a scene-detect pass, and you want them inside the MP4 so the player shows a chapter list. Almost every guide on this shows the same three-line remux command and stops, which leaves the real work undone: producing a correct FFMETADATA file. That file is where chapters silently break.
Quick answer: To make FFmpeg add chapters to an MP4, write a plain text file that starts with;FFMETADATA1and holds one[CHAPTER]block per chapter (TIMEBASE=1/1000,START,END,title), then remux withffmpeg -i in.mp4 -i chapters.txt -map_metadata 0 -map_chapters 1 -c copy out.mp4. Nothing is re-encoded, so a feature-length file finishes in seconds and the video quality is untouched. If you'd rather not host an encoder just to write metadata, FFmpeg Micro runs the same chapter write and remux as one job you call over HTTP: https://www.ffmpeg-micro.com/docs
Chapters are metadata, so the remux costs almost nothing
Adding chapters never needs a re-encode, because chapters live in the container alongside tags like title and artist, not inside the H.264 bitstream. The packets get copied through byte for byte and FFmpeg adds a small text track, so the work is I/O bound. A 45-minute 1080p MP4 goes through in a second or two on a laptop, and the output is the same size plus a few kilobytes.
ffmpeg -i podcast.mp4 -i chapters.txt \
-map_metadata 0 -map_chapters 1 \
-c copy podcast-chaptered.mp4
The -map_metadata 0 is the part most copy-paste answers get wrong. The popular version of this command uses -map_metadata 1, which pulls global metadata from the metadata file and therefore wipes the title, artist, comment, and encoder tags the original file already had. Keep global metadata on input 0 and take only chapters from input 1. Because streamcopy and filtering can't be mixed in one pass, this stays a pure copy job, which is exactly what you want (see why streamcopy and filters don't coexist).
Check the result with ffprobe rather than opening a player:
ffprobe -v error -print_format json -show_chapters podcast-chaptered.mp4
{ "chapters": [
{ "id": 0, "start_time": "0.000000", "end_time": "125.000000",
"tags": { "title": "Cold open" } },
{ "id": 1, "start_time": "125.000000", "end_time": "612.500000",
"tags": { "title": "Why the migration failed" } }
] }
What goes inside an FFMETADATA file
An FFMETADATA file is UTF-8 text with a required first line of ;FFMETADATA1, optional global key/value tags, and one [CHAPTER] section per chapter. Lines beginning with ; or # are comments. FFmpeg reads it through the ffmetadata demuxer, which is why it gets passed as a second input rather than as an option.
;FFMETADATA1
title=Episode 41: The Migration That Ate a Weekend
artist=Ship It Weekly
[CHAPTER]
TIMEBASE=1/1000
START=0
END=125000
title=Cold open
[CHAPTER]
TIMEBASE=1/1000
START=125000
END=612500
title=Why the migration failed
[CHAPTER]
TIMEBASE=1/1000
START=612500
END=3187000
title=Rollback, step by step
Each chapter's START should equal the previous chapter's END. Gaps and overlaps don't throw an error, they just produce a chapter list that doesn't match what the viewer hears. You can also dump an existing file's chapters back into this format with ffmpeg -i input.mkv -f ffmetadata chapters.txt, edit the titles, and remux, which beats hand-typing timestamps you already have.
TIMEBASE decides where your chapters actually land
TIMEBASE is the unit that START and END are counted in, and it's the single most common reason chapters appear in the wrong place with no warning from FFmpeg. With TIMEBASE=1/1000 the values are milliseconds; with TIMEBASE=1/1 they're whole seconds. Transcripts and scene detectors emit seconds, so writing START=125 under a millisecond timebase puts that chapter at 0.125 seconds instead of 2:05, and every chapter in a long episode piles up inside the first second.
| Value written | Under `TIMEBASE=1/1000` | Under `TIMEBASE=1/1` |
|---|---|---|
| `START=125` | 0.125 s | 2:05 |
| `START=612` | 0.612 s | 10:12 |
| `START=612500` | 10:12.5 | 7 days |
Pick milliseconds and convert once at generation time. Second-level granularity forces you to round, and rounding is how a chapter ends up one frame before the cut instead of on it. Don't rely on the demuxer's default unit either, write TIMEBASE explicitly in every block so the file means the same thing to whoever reads it next.
Generating the file from a transcript or a scene pass
Writing the FFMETADATA file takes more work than the remux does. Two sources cover most pipelines: transcript segments from Whisper (or any speech-to-text step that returns start times), and shot boundaries from FFmpeg's own detectors. For the second one, blackdetect gives you the fade-to-black points in seconds:
ffmpeg -i podcast.mp4 -vf "blackdetect=d=0.5:pic_th=0.98" -an -f null - 2>&1 \
| grep black_start
Feed either source into a generator that owns the three things that break: unit conversion, escaping, and the final END.
import re, subprocess
ESCAPE = re.compile(r"([=;#\\])")
def esc(title):
return ESCAPE.sub(r"\\\1", " ".join(title.split()))
def duration_ms(path):
out = subprocess.check_output([
"ffprobe", "-v", "error",
"-show_entries", "format=duration",
"-of", "default=nw=1:nk=1", path,
])
return round(float(out) * 1000)
def write_ffmetadata(video, marks, out_path="chapters.txt"):
# marks: [(start_seconds, title), ...] in ascending order
starts = [round(s * 1000) for s, _ in marks]
ends = starts[1:] + [duration_ms(video)]
lines = [";FFMETADATA1"]
for (start, end), (_, title) in zip(zip(starts, ends), marks):
if end - start < 1000: # skip sub-second chapters
continue
lines += ["", "[CHAPTER]", "TIMEBASE=1/1000",
f"START={start}", f"END={end}", f"title={esc(title)}"]
with open(out_path, "w", encoding="utf-8") as f:
f.write("\n".join(lines) + "\n")
write_ffmetadata("podcast.mp4", [
(0.0, "Cold open"),
(125.0, "Why the migration failed"),
(612.5, "Rollback, step by step"),
])
Whisper's verbose JSON gives you segments[i]["start"] in seconds, so a chapter list is usually a filter over segments where a heading-like phrase appears. If your timestamps arrive as subtitle cues instead, convert them with the same care you'd use for SRT and VTT timestamps: an HH:MM:SS,mmm string has to become an integer, not a float you hope rounds correctly.
Titles containing = ; or # will lose text
FFmpeg splits each metadata line at the first unescaped =, so a title like Q=A with the on-call team is read as the key Q with the value A with the on-call team, and your chapter title disappears with no error. Four characters need a leading backslash inside values:
=because it separates key from value;and#because they start comments\itself, so escape it first or use one regex pass like the code above
Escaped, that title is written title=Q\=A with the on-call team. Non-ASCII titles are fine as long as the file is written as UTF-8, which is why the Python above opens it with an explicit encoding instead of trusting the platform default.
The last chapter's END has to be the real duration
The final chapter is the one players drop. If its END runs past the file's actual duration, or stops well short of it, some players ignore that chapter entirely and others show a chapter list that ends early. Read the duration from the file rather than from whatever your transcript thought it was:
ffprobe -v error -show_entries format=duration -of default=nw=1:nk=1 podcast.mp4
That prints something like 3187.234000, which is 3187234 ms. Transcript segments routinely end 10 to 30 seconds before the real end of the file because the last stretch is music or silence, so using the last segment's end time as the last chapter's END is a reliable way to lose it.
YouTube chapters come from the description, not from the file
YouTube builds chapters from timestamps you put in the video description, so an embedded chapter track does nothing for a YouTube upload. Its rules are specific: the first timestamp must be 0:00, there must be at least three timestamps in ascending order, and each chapter has to be at least 10 seconds long. If you're going from podcast audio to a YouTube video, generate both artifacts from the same marks list, the FFMETADATA file for the downloadable MP4 and a plain text block for the description. The same pipeline that turns the audio into an uploadable video is covered in MP3 to MP4 with a cover image.
Embedded chapters do work in VLC, QuickTime, IINA, most podcast apps, and anything built on libav. MKV stores them natively, MP4 stores them as a text track, and an MP3 output gets ID3v2 chapter frames from the same file.
Pitfalls worth knowing before you run it at scale
Most chapter bugs in production come from the remux, not the metadata. Four to watch for.
A file that already has chapters keeps them unless you say otherwise, so be explicit: -map_chapters 1 replaces them with your file's, and -map_chapters -1 strips chapters entirely. The chapter track in an MP4 shows up as an extra text stream, and downstream tools that don't expect one can fail the way they do on subtitle tracks in MP4 containers. Adding chapters rewrites the moov atom, so if the source was prepared for progressive download, add -movflags +faststart back on this pass. And if you're already re-encoding for delivery, don't run a separate chapter pass, add -i chapters.txt -map_chapters 1 to the encode command you were going to run anyway.
Skip embedded chapters when the destination reads its own format. Instagram, TikTok, and YouTube ignore container chapters, and a podcast host that expects a Podcasting 2.0 chapters JSON file wants that file, not an ID3 frame.
Chaptering inside an automation pipeline
Inside n8n, Make, or Zapier, the FFmpeg part of this is the easy half. The hard half is having a machine with FFmpeg on it, reachable from the workflow, that can pull a 1.2 GB source file and push the result back without timing out the run. FFmpeg Micro does the metadata write and the remux in one job you submit over HTTP and collect by webhook, with no encoder to host and a free tier to test on. The docs have the job shape, and the same call works from an AI agent through the MCP server.
FAQ
How do I add chapters to an MP4 without re-encoding?
Adding chapters to an MP4 without re-encoding takes two inputs and a copy: ffmpeg -i in.mp4 -i chapters.txt -map_metadata 0 -map_chapters 1 -c copy out.mp4. Chapters are container metadata, so the audio and video streams are copied untouched and quality is identical to the source.
What does TIMEBASE mean in an FFMETADATA file?
TIMEBASE is the time unit for a chapter's START and END values. TIMEBASE=1/1000 means the numbers are milliseconds and TIMEBASE=1/1 means whole seconds, so the same value of 125 is 0.125 seconds under the first and 2 minutes 5 seconds under the second.
Why is my last chapter missing?
A missing final chapter almost always means its END value doesn't match the file's real duration. Read the duration with ffprobe -show_entries format=duration and use that number in milliseconds as the last chapter's END instead of the last transcript segment's end time.
How do I read the chapters back out of a video file?
Run ffprobe -v error -print_format json -show_chapters file.mp4 to list every chapter with its start time, end time, and title. To get an editable copy of them, run ffmpeg -i file.mp4 -f ffmetadata chapters.txt, which writes an FFMETADATA file you can change and remux back in.
Can FFmpeg add chapters to MP3 and MKV too?
FFmpeg writes chapters from the same FFMETADATA file to MKV as native Matroska chapters and to MP3 as ID3v2 chapter frames. Only the output extension changes; the metadata file and the -map_chapters 1 flag stay the same.
If your chapter marks already come out of a transcription or scene-detection step, the remux is the last piece of glue you have to own, and it's the piece that wants a server you don't want to run. Sign up free and hand that step an HTTP call instead.
About Javid Jamae
Founder & CEO at FFmpeg Micro
Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.
You might also like

n8n Disabled the Execute Command Node. Your FFmpeg Needs HTTP.
n8n 2.0 disables the Execute Command node by default, breaking FFmpeg workflows on upgrade. How to diagnose it, re-enable it, and replace it for good.

FFmpeg cropdetect isn't one-shot. Detect black bars per file
ffmpeg cropdetect finds black bars on one file fine, then breaks in a batch. Sample past the intro, take the last crop line, round to even, crop per file.

Convert Animated WebP to MP4: FFmpeg 9 Finally Decodes It
FFmpeg 9.0 decodes animated WebP directly, so animated webp to mp4 is now one command. The alpha-channel gotcha, variable frame timing, and the API fallback.
Ready to process videos at scale?
Start using FFmpeg Micro's simple API today. No infrastructure required.
Get Started Free