Skip to content

About

Stop your video automations from dying when n8n can't install ffmpeg. Replace Execute Command ffmpeg steps with one HTTP render call: 9:16 reframe, logo overlay, EBU R128 loudness. MIT, by Avalux.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

n8n-ffmpeg-http-render-kit

Stop your video automations from dying when n8n cannot install ffmpeg. This kit moves the render step out of the n8n container and puts it behind a single HTTP call, so the same workflow produces the same 9:16 clip whether you are on self-hosted n8n, queue mode with five workers, or n8n Cloud. Maintained by Avalux, an AI automation agency that builds and runs production automation for small and mid-sized businesses. Source lives at github.com/elikem2021/avalux-open-source.

Why this exists

Most n8n video workflows start the same way. Someone finds a tutorial that says "add an Execute Command node and run ffmpeg," it works on their laptop, and it ships. Then one of these happens.

The node is not there. n8n Cloud does not expose the Execute Command node. There is no shell you own, no filesystem you control, and no way to install a binary. Any workflow built around shelling out stops dead at the boundary of your own machine, which is usually discovered the week a client asks you to hand the workflow over.

The container refuses the install. The official n8n image is Alpine based and runs as a non-root user. Running apk add ffmpeg inside a live container returns a permission error. The docker exec --user root workaround survives exactly until the next image pull. Teams end up maintaining a forked Dockerfile whose only job is to add one binary, then forget it exists until an upgrade breaks and nobody knows why.

Queue mode multiplies it. In queue mode the render can land on any worker. The binary has to exist, at the same version, on every one of them. One worker built from a slightly older base image produces slightly different output, and you get a client asking why Tuesday's clips sound quieter than Monday's.

Binary data eats memory. Pulling a 300 MB source file into a workflow as binary data means that file crosses node boundaries inside a Node.js process that was not sized for it. Setting N8N_DEFAULT_BINARY_DATA_MODE=filesystem helps, but you are still moving large buffers through an orchestrator whose actual job is coordination.

ffmpeg competes with n8n for CPU. Even when the shell-out works, a four minute render pins cores on the same box that has to poll the queue, answer webhooks, and keep the editor responsive. Executions time out at the default limits, retries stack, and the render restarts from zero.

This kit takes the opposite approach. ffmpeg lives in a small service built for exactly that. n8n makes one HTTP request, gets a job id back, and either polls or waits for a signed callback. Nothing in the workflow depends on what is installed inside the n8n container, which means the workflow is portable between self-hosted and Cloud without edits.

Who this is for

  • Automation builders who render client video inside n8n and keep hitting the Execute Command wall
  • Content and social agencies cutting long-form footage into vertical clips at volume
  • Teams migrating from self-hosted n8n to Cloud, or the other way, who need workflows that survive the move
  • Anyone running n8n in queue mode who is tired of ffmpeg version drift between workers

If you cut fewer than about ten clips a week, you do not need this. A desktop editor is cheaper and better. This starts paying off when the same three transforms run dozens of times a day against footage you did not shoot.

What is in the box

n8n-ffmpeg-http-render-kit/
├── src/
│   ├── server.js          # HTTP surface, job lifecycle, signed callbacks
│   ├── filtergraph.js     # 9:16 reframe, blur pad, logo overlay, presets
│   └── storage.js         # source fetch + output upload (local disk stub)
├── workflows/
│   ├── render-clip-callback.json   # Wait node + $execution.resumeUrl pattern
│   └── render-clip-poll.json       # polling variant for short renders
├── docker/
│   ├── Dockerfile         # node:22-slim + ffmpeg, non-root runtime user
│   └── docker-compose.yml # render service, mounted output volume
├── examples/
│   ├── request-9x16.json
│   └── callback-payload.json
├── .env.example
├── LICENSE
└── README.md

The two code paths that matter are src/filtergraph.js (what ffmpeg is actually told to do) and src/server.js (how a long render is exposed as something n8n can wait on without timing out). src/storage.js is a deliberate stub. Swap it for S3, R2, GCS, or whatever your clients already pay for.

Quick start

git clone https://github.com/elikem2021/n8n-ffmpeg-http-render-kit.git
cd n8n-ffmpeg-http-render-kit
cp .env.example .env

# generate the two secrets the service needs
openssl rand -hex 32   # -> RENDER_API_TOKEN
openssl rand -hex 32   # -> RENDER_CALLBACK_SECRET

docker compose -f docker/docker-compose.yml up --build

Confirm ffmpeg is where the service expects it:

docker compose -f docker/docker-compose.yml exec render ffmpeg -version

Fire a render:

curl -X POST http://localhost:8080/v1/renders \
  -H "Authorization: Bearer $RENDER_API_TOKEN" \
  -H "Idempotency-Key: demo-001" \
  -H "Content-Type: application/json" \
  -d @examples/request-9x16.json

You get a 202 with a job id. Poll it:

curl -s http://localhost:8080/v1/renders/<id> \
  -H "Authorization: Bearer $RENDER_API_TOKEN" | jq

Then import workflows/render-clip-callback.json into n8n and point the HTTP Request node at your service URL.

The render contract

One endpoint does the work. Everything else is status.

POST /v1/renders

{
  "source": { "url": "https://cdn.example.com/raw/interview-04.mp4" },
  "preset": "reel_1080x1920",
  "overrides": {
    "anchor": 0.42,
    "logo": { "url": "https://cdn.example.com/brand/mark.png", "widthPct": 0.12, "position": "top-right", "marginPx": 64 },
    "loudness": { "i": -14, "tp": -1.5, "lra": 11 },
    "trim": { "startSec": 122.5, "durationSec": 48 }
  },
  "callback": { "url": "https://n8n.example.com/webhook-waiting/abc123" }
}

Idempotency-Key is required. n8n retries. If a workflow re-fires the same request after a network blip, you want the same job id back, not a second render billed to the same client. The reference implementation keys the map on that header.

The response is 202 Accepted with { id, status, statusUrl }. Status moves through queued, probing, rendering, then succeeded or failed. On success the job carries the output URL, the final duration, dimensions, and the measured integrated loudness so you have something to assert on in a QA step.

How the 9:16 reframe is built

The interesting part is what you lose. Take a normal 1920x1080 source and crop it to 9:16. The crop width is ih * 9/16, which is 1080 * 0.5625 = 607.5, rounded down to an even 608 pixels. That is 31.7 percent of the original frame width. Two thirds of the shot is gone, and whatever was framed at the edges goes with it. Then that 608x1080 region gets scaled up to 1080x1920, a 1.78x upscale, which is where soft footage starts to look obviously soft.

So the kit ships two fit modes and makes you pick.

crop keeps full resolution in the kept region and accepts the loss. The crop window is not locked to center. The anchor value moves it horizontally from 0.0 (hard left) to 1.0 (hard right):

crop=w=ih*9/16:h=ih:x=(iw-ih*9/16)*0.42:y=0

For a talking-head interview shot off-center, an anchor of about 0.35 to 0.45 usually keeps the subject in frame where a center crop would slice them. This is the single most useful parameter in the whole kit and the one worth wiring to a per-client field in your CRM or sheet.

blur_pad keeps the entire frame. It splits the input, scales one copy to cover 1080x1920 and blurs it as a background plate, scales the other to fit inside, and centers it:

[0:v]split=2[bg][fg];
[bg]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,gblur=sigma=22[bgb];
[fg]scale=1080:1920:force_original_aspect_ratio=decrease[fgs];
[bgb][fgs]overlay=(W-w)/2:(H-h)/2[v]

gblur is slower than boxblur and looks better. On a four core box the blur plate is usually the dominant cost of the render, so if throughput matters more than polish, boxblur=10:2 is the trade.

Both modes end with setsar=1 and fps normalization. Mixed frame rate sources are the usual cause of clips that play fine locally and stutter after a platform re-encode.

Loudness normalization that matches what platforms do

The common mistake is a single-pass loudnorm. In one pass the filter works dynamically without knowing the program's real measurements, which on speech with music beds produces audible pumping. This kit runs two passes.

Pass one measures and prints JSON:

ffmpeg -i in.mp4 -af loudnorm=I=-14:TP=-1.5:LRA=11:print_format=json -f null -

That returns input_i, input_tp, input_lra, input_thresh, and target_offset. Pass two feeds them back in with linear=true, which applies a single gain instead of riding the level:

-af loudnorm=I=-14:TP=-1.5:LRA=11:measured_I=-21.4:measured_TP=-3.1:measured_LRA=6.2:measured_thresh=-32.7:offset=-0.3:linear=true:print_format=summary

If the measured LRA is wider than the target, ffmpeg falls back to dynamic mode and says so in the summary. The service captures that and surfaces it as a dynamicFallback flag on the job, because it is a real quality signal and you want it in your logs rather than discovered by a client.

One detail that costs people hours: loudnorm resamples internally to 192 kHz, and if you do not pin -ar 48000 on the output you can end up with an audio stream that some players and some upload pipelines reject. The default target of -14 LUFS integrated with -1.5 dBTP is the common rule of thumb for social platforms that normalize playback. It is a starting point, not gospel. Broadcast delivery is a different target and this kit is not the right tool for it.

Logo overlay and platform safe areas

The logo is a second input scaled relative to the output width rather than to a fixed pixel size, so one brand asset works across presets:

[1:v]scale=w=iw*0:h=-1[lg]   # widthPct resolves to round(1080*0.12)=130 before the graph is built
[v][lg]overlay=x=W-w-64:y=64

Placement matters more than size. On a 1080x1920 vertical clip, the platform's own interface eats real estate: roughly the bottom 300 to 350 pixels for caption and handle text, and roughly 140 to 160 pixels along the right edge for the action rail. Anything you burn into those regions competes with the UI and often ends up half covered. The presets keep overlays out of a bottom safe band of 320 pixels and a right band of 150 pixels by default, and position accepts top-left, top-right, bottom-left, and bottom-right with the margin applied from the safe band rather than the frame edge.

If you burn in captions, do it inside the same render call rather than as a second pass. Every additional encode generation is another round of quality loss on footage that is often already compressed.

Async by default, because the HTTP Request node will time out

A 90 second clip with a blur plate and two passes of loudness analysis is not a request you hold open. The kit is async, and the n8n side uses the pattern n8n already has for this.

In workflows/render-clip-callback.json, the HTTP Request node posts the job and passes {{ $execution.resumeUrl }} as callback.url. The next node is a Wait node set to resume on webhook call. The execution parks, costs nothing while parked, and wakes when the render service posts the finished payload back to that exact URL. No polling loop, no arbitrary Wait node with a guessed duration, no execution timeout at minute five.

Callbacks are signed. The service sends X-Avalux-Timestamp and X-Avalux-Signature: sha256=<hex> over timestamp + "." + rawBody using RENDER_CALLBACK_SECRET. Verify it in a Code node before you trust the payload, and reject timestamps older than five minutes. A resume URL is effectively a public endpoint the moment it leaves your network.

For renders under about 20 seconds, workflows/render-clip-poll.json is simpler: post, wait 5 seconds, GET status, loop with a bounded counter. Use it for thumbnails and trims, not for anything with a blur plate.

What this kit deliberately does not do

This is scaffolding, not a product. Running it for real clients means adding the parts that are specific to you:

  • The job store is an in-memory Map. Restart the process and every in-flight job vanishes. Redis or Postgres before anything real.
  • There is no queue or concurrency limit. Ten simultaneous requests will start ten ffmpeg processes and take the box down. BullMQ or a simple semaphore, plus a worker count tied to actual cores.
  • Storage is a local disk stub. No signed URLs, no lifecycle expiry, no per-client bucket prefixes.
  • Auth is one shared bearer token. Fine for a single internal service, wrong the moment two clients share an instance.
  • Source URLs are not validated. Fetching arbitrary user-supplied URLs server-side is an SSRF vector. Allowlist your own domains.
  • No observability. No metrics, no per-job cost accounting, no alerting when the dynamic loudness fallback fires on a whole batch.
  • Presets are opinionated guesses. The CRF, preset speed, and loudness targets should be tuned against your actual footage and your actual platforms.

None of that is hidden work. It is just work, and it is the difference between a repo that runs and a pipeline you can bill against.

Why we built this

We kept inheriting the same broken workflow. A client would come to Avalux with an n8n automation that cut social clips, it worked for a month, and then it stopped, usually right after an n8n upgrade or a move to Cloud. Every time, the root cause was the same: a render step that assumed it could reach a shell. Rebuilding that boundary correctly takes a day of work that nobody wants to pay for twice, so we published the boundary.

What is here is the shape of the fix, not the finished pipeline. If you want the finished pipeline, with the queue, the storage, the per-client presets, and the monitoring that tells you a batch went quiet before the client does, that is the work we do. Reach us at avalux.io.

Avalux's other open source projects

  • freight-eta-toolkit - ETA math, geofencing, and carrier API normalization for freight and logistics ops
  • avalux-open-source - index of everything we have published, with notes on what each kit is and is not

Issues and pull requests are welcome on all of them. If you fix something here that bit you in production, that is the most useful kind of contribution.

License

MIT. See LICENSE. Use it commercially, fork it, ship it inside client work, no attribution required. If it saves you a day, a star helps other people find it.

Keywords

n8n ffmpeg not working, n8n execute command disabled, n8n video rendering without ffmpeg, n8n ffmpeg docker install, n8n cloud execute command node missing, apk add ffmpeg permission denied n8n, n8n video automation, n8n 9:16 vertical video, ffmpeg http render api, n8n render service, n8n binary data large video, n8n queue mode ffmpeg, loudnorm two pass n8n, ebu r128 normalization api, ffmpeg logo overlay automation, n8n social clip automation, n8n wait node resume url, self-hosted n8n video pipeline

About

Stop your video automations from dying when n8n can't install ffmpeg. Replace Execute Command ffmpeg steps with one HTTP render call: 9:16 reframe, logo overlay, EBU R128 loudness. MIT, by Avalux.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages