I had a handful of long travel videos already published and wanted a steady cadence of vertical Shorts without generating AI footage or re-uploading old clips. So I asked an AI coding agent to build a small local ffmpeg pipeline: segment lists in, Shorts-ready MP4s out, with every cut checked against what is actually on camera.
Everything runs on a local machine with ffmpeg and Python. No cloud video editor, no upload until a human has approved the stills and titles.
What the agent built
- Two small scripts:
build_short.py(one Short from a list of segments) and a batch script that builds a week’s worth - Three vertical Shorts cut from downloads of existing long videos. These are real footage, not AI footage or re-uploads of old Shorts.
- Output at 1080×1920, 30 fps, H.264 High at CRF 18, AAC 192k / 48 kHz, with faststart for streaming
- A hook caption on the first ~2 seconds (bold white text with a dark outline, faded out)
- Audio kept in sync with 40 ms fades at each cut and a short fade-out at the end
The AI part that actually mattered: checking what’s on camera
The original plan had a topic for each Short, written from memory. Before cutting anything, the agent generated contact sheets from each long video and compared them with the plan:
# One frame every 10 s, tiled into a single image the agent can review
ffmpeg -i long.mp4 -vf "fps=1/10,scale=320:-1,tile=6x5" -frames:v 1 sheet-01.png
Several planned topics weren’t in the footage at all. The agent flagged them instead of inventing a matching filename and caption. I renamed the Shorts to match what is visible, and the missing topics went on a list for future filming. That one check is the most useful thing the agent did in this project. It stopped me from publishing titles that the video doesn’t back up.
Crop vs blur-fill (honest tradeoffs)
The sources are 720p. A 9:16 centre crop of a 1280×720 frame is about 405 px wide, which then gets upscaled roughly 2.67× to 1080. That works, but it’s softer than footage shot vertically.
# Centre crop to 9:16
ffmpeg -ss 00:01:12 -to 00:01:19 -i long.mp4 \
-vf "crop=ih*9/16:ih,scale=1080:1920:flags=lanczos" seg-a.mp4
For wide subjects the agent used blur-fill: a blurred full-frame background with the sharp frame letterboxed on top, so the whole subject stays in shot instead of being cropped off.
ffmpeg -i seg.mp4 -filter_complex \
"[0:v]split=2[bg][fg]; \
[bg]scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,boxblur=20:2[bgb]; \
[fg]scale=1080:-2[fgs]; \
[bgb][fgs]overlay=(W-w)/2:(H-h)/2" seg-b.mp4
Centre crop stays the default for tight shots. Blur-fill is per segment and switched on in the segment list.
One encode profile, then verify
ffmpeg -i joined.mp4 -c:v libx264 -profile:v high -crf 18 -r 30 -pix_fmt yuv420p \
-c:a aac -b:a 192k -ar 48000 -movflags +faststart short-01.mp4
ffprobe -v error -show_entries format=duration:stream=codec_name,width,height,r_frame_rate short-01.mp4
The agent ran ffprobe on every output to confirm duration, resolution, frame rate, and that an audio stream exists before anything was handed to me.
What I refused to invent
- No filename or caption for a subject that isn’t on camera
- No AI-generated B-roll to fill gaps in the plan
- No subscriber, view, or “this will beat the median” forecasts
- No upload until the stills and titles were approved
Ship checklist (if you copy the pattern)
- Download the long video and have the agent build contact sheets before anyone names a Short.
- Keep one small script that takes a segment list (start, end, crop or blur-fill, crop centre).
- Encode once to a Shorts-safe profile (1080×1920, 30 fps, H.264 High, AAC) and check every file with
ffprobe. - Name each Short after what’s visible, and leave missing topics for a future shoot.
- Get a human to approve stills and titles before uploading or scheduling.
FAQ
Why not AI-generated B-roll?
I already had real footage. Cutting what exists is easier to trust than synthesizing shots that never happened.
Does this need a GPU?
No. These are short clips and CPU x264 encoding is fine. A GPU encoder would be faster but isn’t needed at this volume.
Why not upload in the same session?
Stills and titles get approved first, and the platform’s copyright check on background audio is still a human pass before anything is scheduled.
Closest related Builders notes?
Whisper.cpp on Ubuntu is where ffmpeg shows up as the boring audio prep step for local transcription, and 7 Rules for Letting an AI Agent Maintain Your Homelab covers the guardrails I use when an agent runs commands for me.
