ffmpeggpu-transcodingvideo-encoding

scale_npp Removed in FFmpeg 9: Move GPU Scaling to scale_cuda

·Javid Jamae·9 min read
scale_npp Removed in FFmpeg 9: Move GPU Scaling to scale_cuda

You upgraded to FFmpeg 9 and your GPU transcode died with Unknown filter 'scale_npp'. Or your build script died earlier than that, on a configure line that has carried --enable-libnpp since 2019. Both failures have the same cause and the same fix.

Quick answer: Seeing scale_npp removed from your FFmpeg build isn't a packaging mistake. FFmpeg 9.0 deleted the entire libnpp filter set, so scale_npp, scale2ref_npp, sharpen_npp, and transpose_npp no longer exist. Replace scale_npp=W:H with scale_cuda=W:H and transpose_npp=clock with the new transpose_cuda=clock; sharpen_npp has no CUDA replacement and now costs a round trip through system memory. If you rented a GPU purely to keep FFmpeg's scaler off the CPU, price that box against sending the same resize to FFmpeg Micro as one API call with no servers to run: ffmpeg-micro.com/pricing.

Why libnpp is gone, and why you may not miss it

FFmpeg 9.0 "Lei" shipped on August 3, 2026, and it removed NVIDIA Performance Primitives support outright rather than deprecating it first. The removal took four filters with it, plus the --enable-libnpp configure option, which is why some people hit this at build time and others hit it at runtime. There's an open issue titled "FFMPEG and NVidia no longer supporting scale_npp" collecting the confused, and the 9.0 release notes are the authoritative answer.

Conventional wisdom says losing scale_npp costs you GPU scaling. Most migration write-ups read exactly that way. But scale_npp was never in a build you could hand to anyone: libnpp is NVIDIA-licensed, so it required --enable-nonfree, which made the resulting binary legally undistributable. That's why Ubuntu's ffmpeg package never had scale_npp, and why you were compiling your own in the first place.

scale_cuda has no such flag. It compiles with the ordinary CUDA path and ships in plenty of stock builds. The filter you lost was the one you had to maintain a private toolchain for.

The rewrite table: every -vf chain, before and after

Most scale_npp chains convert by changing one word, and the ones that don't are the ones with a non-scale NPP filter in them. Run ffmpeg -filters | grep cuda on your new binary first so you know what your build actually has.

FFmpeg 8.x chainFFmpeg 9.0 chainWhat changes
`scale_npp=1280:720``scale_cuda=1280:720`Filter name only
`scale_npp=1920:-1``scale_cuda=1920:-2`Same auto-height math; `-2` keeps it even for H.264
`scale_npp=1280:720:interp_algo=super``scale_cuda=1280:720:interp_algo=lanczos``super` doesn't exist in scale_cuda
`scale_npp=1280:720:format=yuv420p``scale_cuda=1280:720:format=yuv420p`Option carries over unchanged
`scale_npp=...:force_original_aspect_ratio=decrease``scale_cuda=...:force_original_aspect_ratio=decrease`Carries over, as does `force_divisible_by`
`transpose_npp=clock``transpose_cuda=clock`New filter in 9.0, same `dir` values
`scale_npp=1280:720,sharpen_npp=...``scale_cuda=1280:720,hwdownload,format=nv12,unsharp=5:5:0.8,hwupload_cuda`No CUDA sharpener exists
`scale2ref_npp`No equivalentProbe the reference stream, pass literal dimensions

A full command, before and after:

# FFmpeg 8.x
ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i in.mp4 \
  -vf scale_npp=1280:720:interp_algo=super \
  -c:v h264_nvenc -b:v 4M out.mp4

# FFmpeg 9.0
ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i in.mp4 \
  -vf scale_cuda=1280:720:interp_algo=lanczos \
  -c:v h264_nvenc -b:v 4M out.mp4

Interpolation names don't map one to one

The two filters use different vocabularies for the same math, so a copied interp_algo value is the most common second failure after the rename. scale_npp accepted nn, linear, cubic, cubic2p_bspline, cubic2p_catmullrom, cubic2p_b05c03, super, and lanczos. scale_cuda accepts nearest, bilinear, bicubic, and lanczos.

So nn becomes nearest, linear becomes bilinear, and anything cubic becomes bicubic. The one real loss is super, NPP's supersampling mode, which was the best option for big downscales like 4K to 720p. Use lanczos in its place and expect slightly different ringing on hard edges. The defaults differ too, so name the algorithm explicitly on both sides instead of trusting whatever the filter picks.

Pixel formats scale_cuda handles now

scale_cuda in FFmpeg 9.0 got a generic filtering path underneath it, which is what unlocked 10-bit and 12-bit 4:2:2 and 4:4:4 scaling on the GPU. That's a capability scale_npp never had, so a few chains that used to force a CPU round trip for high bit depth chroma can now stay in VRAM.

NV12 and P010 still work the way they always did. If you're feeding 10-bit 4:2:2 from a camera or a ProRes intermediate, check ffmpeg -h filter=scale_cuda on your build for the exact format list rather than assuming, because this is the part of the filter that changed most between point releases.

The round trip you were avoiding just moved

Running scale_npp was about keeping frames in GPU memory from decode to encode. That still holds for pure scaling, but sharpen_npp has no CUDA-native replacement in FFmpeg 9.0, so sharpening a GPU-decoded frame now costs a hwdownload to system memory and an hwupload_cuda back. Every frame, both directions, over PCIe.

ffmpeg -hwaccel cuda -hwaccel_output_format cuda -i in.mp4 \
  -vf "scale_cuda=1280:720,hwdownload,format=nv12,unsharp=5:5:0.8,hwupload_cuda" \
  -c:v h264_nvenc -b:v 4M out.mp4

If your build has OpenCL, unsharp_opencl with CUDA-OpenCL interop avoids the host bounce, at the cost of a second hardware context to configure and debug. For most pipelines that's more operational surface than the sharpening is worth.

The same applies to any CPU filter you drop into a CUDA chain. crop, drawtext, and overlay can't read CUDA frames, and FFmpeg will tell you so rather than insert the conversion for you.

Common pitfalls when moving to scale_cuda

These four account for nearly every migration that fails after the filter rename is done.

  1. Missing -hwaccel_output_format cuda. With only -hwaccel cuda, frames land back in system memory and scale_cuda can't accept them. You get Impossible to convert between the formats supported by the filter 'graph 0 input from stream 0:0' and the filter 'auto_scale_0', which reads like a pixel format problem and isn't.
  2. hwdownload with no format after it. hwdownload alone leaves an ambiguous format and errors out. Always follow it with format=nv12 (or format=p010le for 10-bit input).
  3. A configure line that still says --enable-libnpp. FFmpeg 9's configure rejects it as an unknown option, so the build fails before any of your filter chains get a chance to.
  4. Assuming scale2ref_npp has a CUDA successor. It doesn't. Read the reference stream's dimensions with ffprobe -v error -select_streams v:0 -show_entries stream=width,height -of csv=p=0 ref.mp4 and pass the numbers into scale_cuda as literals.

If your scripts are failing in more places than this one filter, FFmpeg 9 also removed -vsync, -top, -qphist, and -filter_complex_script, which we walked through in FFmpeg 9 breaking changes, and it flipped TLS verification on by default, covered in FFmpeg 9.0 verifies TLS.

What the rented GPU costs you now

A broken filter chain is a good moment to ask whether the rented GPU is still earning its keep. Public price comparisons for RunPod and Vast.ai put an RTX 4090 around $0.32 to $0.34 an hour. At 730 hours a month that's roughly $234 to $248 for a machine that's always on, and first-hand accounts of running transcodes on unverified Vast.ai hosts report effective costs 20% to 40% above sticker once instance restarts and dead time are counted. Call it $280 to $350 a month, before storage and egress.

Now count the actual work. A 1080p to 720p NVENC transcode runs far faster than realtime, so a pipeline doing a few thousand short clips a month is buying hours of idle GPU for minutes of scaling. The GPU made sense when it was the only way to avoid a CPU-bound swscale bottleneck at volume. It makes much less sense as a permanently provisioned answer to a bursty queue.

That's the version of this problem FFmpeg Micro solves: the same resize, crop, transpose, or format conversion goes out as one API call, you get a webhook or a polled job result back, and there's no encoder host to patch through the next FFmpeg major. There's a free tier to start and usage-based pricing after, which is a different shape of bill than a 24/7 instance.

When keeping the GPU is still right

Sustained load is the honest case for staying on your own hardware. If you're transcoding live streams continuously, or pushing thousands of hours of video a day, a GPU that never idles beats per-job pricing and you should finish the scale_cuda migration and move on.

Archival-quality downscales are the other one. CPU scale with flags=lanczos through swscale still produces cleaner results than any GPU scaler, and for a mastering step where the encode takes minutes anyway, the GPU saves you very little. Keep a CPU path for those and let the GPU handle the volume work.

FAQ

Is scale_npp coming back in a later FFmpeg 9 release?

scale_npp is not coming back. FFmpeg removed libnpp support rather than deprecating it, and neither 9.0.1 (August 12, 2026) nor 9.0.2 (September 18, 2026) restored it. Plan the scale_cuda rewrite instead of waiting for a point release.

Can I keep using scale_npp by staying on FFmpeg 8?

Staying on FFmpeg 8 does keep scale_npp working. The 8.1 branch is still maintained and 8.1.3 shipped on September 21, 2026 with backported fixes. That buys time, not a permanent answer, and you're still on a nonfree build you can't redistribute.

Is scale_cuda slower than scale_npp?

For a plain resize, the difference between scale_cuda and scale_npp is small compared to the cost of decode and encode around it, and both keep frames in VRAM. The throughput regression people actually notice comes from chains that used sharpen_npp, because those now bounce through system memory.

What replaces transpose_npp in FFmpeg 9?

transpose_cuda replaces transpose_npp and is new in FFmpeg 9.0. It takes the same dir values, so transpose_npp=clock becomes transpose_cuda=clock with no other changes to the chain.

Does scale_cuda need a special FFmpeg build?

scale_cuda needs CUDA support compiled in, usually via --enable-cuda-nvcc with the ffnvcodec headers installed, but it does not need --enable-nonfree the way scale_npp did. Check with ffmpeg -filters | grep scale_cuda before you rewrite anything.

If the GPU was only ever there to run one filter, the cheaper migration is not rewriting the chain at all. Sign up free and send one resize through the API to see what your queue actually costs without a box behind it.

About Javid Jamae

Founder & CEO at FFmpeg Micro

Javid is a software engineer, author, and entrepreneur with over 25 years of professional software development experience across enterprise, startup, and consulting environments. He founded FFmpeg Micro to make video processing accessible to developers through a simple, automation-first REST API.

Software EngineeringVideo ProcessingFFmpegCloud ArchitectureAPI DesignAutomation

Ready to process videos at scale?

Start using FFmpeg Micro's simple API today. No infrastructure required.

Get Started Free