mcpai-agentsapi

The OpenAI Agents API can't cut video. MCP gives it FFmpeg.

··10 min read
The OpenAI Agents API can't cut video. MCP gives it FFmpeg.

Your agent can read a transcript, pick the three best moments, and write the hook text for each one. Then it hands you timestamps and stops, because trimming a clip, burning captions, and reframing to 9:16 are not things a language model does. The missing piece isn't reasoning. It's that there's no FFmpeg binary anywhere in the agent's runtime.

Quick answer: To give an OpenAI agent video processing tools, register an MCP server as a tool in the request. OpenAI Agents API MCP support treats MCP servers as a first-class tool source, so the agent discovers the server's tools at runtime. The DIY route is hosting FFmpeg behind your own HTTP service and hand-writing a tool schema for every operation you want. Pointing the agent at https://mcp.ffmpeg-micro.com instead gives it trim, caption, watermark, and reframe as one tool call, with no FFmpeg build to maintain and no server to run.

Conventional wisdom says the hard part of an agentic video pipeline is the model. Most first attempts do look that way: the agent hallucinates a timestamp, the clip lands a second late, and you go tune the prompt. But the failure usually isn't the model's judgment, it's the shape of the tool underneath it. Video work is slow, stateful, and file-based, and a tool that pretends otherwise breaks the agent loop in ways no prompt fixes.

An agent with no media tools can only produce instructions

An OpenAI Agents API agent runs in a managed runtime with session orchestration, context compaction, and recovery, and it can run in an OpenAI-hosted sandbox or on your own infrastructure. None of that includes a media encoder. Ask it to "cut 0:42 to 1:04 and burn the captions" and the best it can do is emit an FFmpeg command as text for a human to run.

The work itself is three ordinary FFmpeg invocations. Trimming a clip:

ffmpeg -ss 00:00:42 -i source.mp4 -t 22 -c:v libx264 -crf 20 -c:a aac clip.mp4

Burning a subtitle file into the picture:

ffmpeg -i clip.mp4 -vf "subtitles=clip.srt:force_style='Fontsize=24,Outline=2'" -c:a copy captioned.mp4

Cropping a 1920x1080 master to a 1080x1920 vertical:

ffmpeg -i captioned.mp4 -vf "crop=ih*9/16:ih,scale=1080:1920" vertical.mp4

Every one of those needs a binary with libx264 and libass compiled in, a filesystem with room for the intermediate files, and minutes of CPU. Agent sandboxes are built for short, cheap tool calls. This is why the pattern that works is not "install FFmpeg next to the agent" but "give the agent a tool that owns the encoder."

Registering an MCP server as a tool source

Adding an MCP server to an agent request takes a single tool entry. You give it a label, the server URL, an auth header if the server needs one, and the list of tools you want exposed.

{
  "model": "gpt-5.1",
  "tools": [
    {
      "type": "mcp",
      "server_label": "ffmpeg_micro",
      "server_url": "https://mcp.ffmpeg-micro.com",
      "headers": { "Authorization": "Bearer $FFMPEG_MICRO_API_KEY" },
      "allowed_tools": ["trim_video", "caption_video", "resize_video"],
      "require_approval": "never"
    }
  ],
  "input": "Cut 00:00:42 to 00:01:04 from https://cdn.example.com/ep12.mp4, burn in captions, and give me a 9:16 version."
}

The agent fetches the server's tool list on the first turn and gets names, descriptions, and JSON schemas without you writing any of them. That's the real reason MCP beats a hand-rolled function tool here: when the server adds a watermark or audio-extraction tool next month, your agent picks it up with no code change. The FFmpeg Micro MCP server exposes the same job catalog the REST API does, so an agent and a cron script hit identical processing, and the docs list the exact tool names and parameters the server advertises.

Use allowed_tools deliberately. A server with twenty tools puts twenty schemas into context on every turn, and an agent that can see delete_job will eventually call it.

The async problem most MCP tutorials skip

A video job takes longer than a tool call should. Transcoding a 10-minute 1080p file runs tens of seconds to a few minutes depending on the codec and filters, and burning captions adds a full re-encode on top. If your tool blocks until the file is done, you are holding an agent turn open for three minutes and praying nothing times out.

The job API is built around that reality. A submit call comes back immediately with an identifier and a status:

{ "id": "job_8Qa2...", "status": "queued" }

The agent's turn ends there. Three ways to pick the result back up, in order of how well they behave:

  1. Webhook into your app, then resume the session. The job posts its completion to your endpoint, your handler writes the output URL into the session, and the agent continues with durable sessions doing the state keeping. No idle tokens, no polling.
  2. One deferred poll. If you can't take an inbound webhook, have the agent call the status tool once after a sensible wait rather than in a loop. A 22-second clip is not ready in 2 seconds, and ten "status":"processing" responses in context make the model start inventing output URLs.
  3. Chain jobs server-side. Trim, then caption, then reframe is one submission with three steps, not three agent turns. The fewer round trips the model mediates, the fewer chances it has to drop a file reference.

Pass URLs, never bytes. Base64-encoding a 40 MB MP4 into a tool argument blows the context window before the job starts. The same lesson applies outside agents entirely: large video transcription in n8n fails for exactly this reason.

Hand-rolled FFmpeg service vs. an MCP tool call

Both paths end with the same MP4. What differs is how much of the media stack you own and how much of it the agent can break.

ConcernFFmpeg behind your own endpointMCP server tool call
EncoderYou build and patch it. FFmpeg 9 removed `-vsync`, `-top`, and every libnpp filterManaged; your agent never sees a flag
Tool schemasOne JSON schema per operation, written and versioned by youDiscovered from the server's tool list
Long jobsYour own queue, workers, and timeout handlingJob id, poll or webhook, built in
Cold startsA container with libx264 and libass is not smallNone
Cost floorAn always-on worker, whether or not jobs arriveFree tier, then usage-based

The FFmpeg 9 point is not hypothetical. Frame-extraction scripts in agent tooling repos broke this summer when -vsync vfr was hard-removed, and the projects hitting it first were LLM pipelines pulling stills for vision models. If you own the binary, you own that migration.

Pitfalls that show up on the first real run

The approval default catches almost everyone. OpenAI's hosted MCP tool asks for human approval before calling a remote tool unless you set require_approval, so an unattended agent sits there waiting for a confirmation nobody is going to give.

  • Trimming with -c:v copy lands on the wrong frame. Stream copy can only cut at keyframes, so a request for 00:00:42 silently becomes 00:00:40. Re-encode when the boundary matters.
  • Captions burned before reframing get cropped off. Crop to 9:16 first, then burn text, or your subtitle line loses its edges. Order of operations is the single most common bad output in agent-built pipelines.
  • The agent reuses a stale output URL. If a job's result URL expires, a retry two turns later fetches nothing. Have the agent re-resolve the job by id instead of caching the link in its reasoning.
  • An agent is the wrong driver for a fixed pipeline. If every input gets trimmed to 30 seconds, captioned, and reframed with no decisions in between, that's a workflow, not an agent. Run it in n8n, Make, or a plain script and spend the model tokens somewhere a judgment call actually happens.

That last one is worth taking seriously. An LLM in the loop earns its keep when it's choosing which 22 seconds, not when it's executing a known recipe. For the known recipes, the same jobs run from n8n over plain HTTP.

The same server works for Claude, Cursor, and n8n agents

MCP is a protocol, not an OpenAI feature, so nothing in this setup is specific to the Agents API. Claude Desktop, Claude Code, Cursor, Zed, and n8n's AI Agent node all take a Streamable HTTP MCP server URL, and Make documents adding MCP servers on every plan including the free tier. Point any of them at the same server URL with the same API key and the tool catalog is identical.

That portability is the practical argument for the MCP route over a custom function tool. A hand-written OpenAI function schema is OpenAI-shaped. A server URL moves with you when you swap models, which agent builders have been doing roughly every quarter.

FAQ

Does the OpenAI Agents API support MCP servers in public beta?

The Agents API entered public beta with MCP servers listed as a supported tool source alongside custom tools, durable sessions, and progress streaming. You add a server by including a tool of type mcp with a server_url in the request, and the agent reads the tool list from the server itself.

How do I stop an agent from blocking on a long video job?

Submit the video job, let the tool return a job id immediately, and end the agent turn. Pick the result up through a webhook that resumes the session, or have the agent call a status tool once after a delay. Polling in a tight loop fills the context with status payloads and pushes the model toward fabricating output URLs.

Can an AI agent do video processing without FFmpeg installed?

An AI agent can do video processing with no local FFmpeg by calling an MCP server that runs the encoder remotely. The agent sends parameters like a start timestamp and a target aspect ratio, and the server returns a finished file, so the agent's sandbox never needs libx264, libass, or disk space for intermediates.

What video jobs are worth exposing to an agent?

Trim, caption, reframe to 9:16 or 1:1, watermark, extract audio, and pull thumbnails cover most of what an agent gets asked for in content pipelines. Aspect-ratio work in particular is unavoidable with AI-generated footage, since Veo and Sora render only 16:9 and 9:16 and every other placement is a post-generation crop.

Is MCP better than a plain HTTP tool for video?

MCP is better when the tool catalog changes or the agent moves between runtimes, because the agent discovers tools at runtime instead of reading a schema you maintain. A plain HTTP call is fine for one fixed operation in one fixed app, and the same backend usually serves both.

Grab an API key on the free tier, drop https://mcp.ffmpeg-micro.com into your agent's tool list, and ask it to cut a clip. The first trim is a one-line config change away.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Ready to process videos at scale?

Start using FFmpeg Micro's simple API today. No infrastructure required.

Get Started Free