Auto captions for reels with an open-source editor means you own the whole caption pipeline: the transcription model, the timing data, the font files and the render script. Nothing is locked behind a subscription, and you can change one word without paying to re-render a whole project. This guide walks through a working local stack for that, then shows how to push the finished reels out to Instagram, TikTok, YouTube and the rest from one place.
raw clip (.mp4)
|
v
[ Whisper ] --> transcript.json (word-level timings)
|
v
[ brand config ] --> font, color, position, safe area
|
v
[ FFmpeg / Remotion ] --> burned-in captions
|
v
finished reel --> Creator OS --> Instagram, TikTok, YouTubeWhy Auto Captions for Reels Matter for Reach and Retention
Most people watch reels with the sound off at least part of the time. If your reel has no captions, those viewers get a person moving their mouth and nothing else. Captions give them a reason to keep watching, and watch time is the signal every platform reads first.
Captions also change how you edit. Once you know every word has to be legible on a phone, you start writing shorter sentences, cutting filler and giving the on-screen text room to breathe. That forces tighter scripts.
There is a second, less obvious benefit. A transcript is text. Text is searchable, reusable and easy to turn into a blog post, a comment reply or a pinned first comment with your links. If you want to see what a finished, agent-driven version of that looks like, the open source Claude video editor walkthrough covers it end to end.
What Open-Source Editors and Tools Can Actually Do
Open-source caption tooling splits into three jobs, and it helps to pick one tool per job instead of looking for a single app that does everything.
- Transcription. Whisper and its faster forks turn audio into text with word-level timestamps. They run locally, they handle accents and background noise reasonably well, and you can pick model sizes based on how much patience you have.
- Burning in. FFmpeg draws text onto video with full control over font, size, outline, position and timing. It is unglamorous and it never surprises you.
- Animated styling. Remotion lets you build caption components in React, so a caption can scale, bounce and change color per word based on the timings in your transcript.
A social media scheduling dashboard handles the publishing side of captions, meaning the caption text under the post. That is a different job from burned-in subtitles. Creator OS splits it the same way: your render happens locally, and Creator OS is where the finished file gets posted, captioned and tracked. If you want a render pipeline that is already wired up as Claude Code skills, look at the open edits update.
Set Up Your Open-Source Caption Stack (Whisper, FFmpeg, Remotion)
Three installs, one folder per project. Keep the folder flat so scripts can find things without configuration.
reel-01/
raw.mp4
transcript.json
brand.json
render.py
out/reel-01-captioned.mp4
Whisper gives you the transcript with timings. FFmpeg does the draw. Remotion handles anything that moves. A brand.json file holds the font path, hex colors, outline width, margin from the bottom edge and the size ratio, so a caption style is data instead of code you edit every time.
{
"font": "Inter-ExtraBold.ttf",
"primary": "#FFFFFF",
"highlight": "#FFD400",
"outline": 6,
"bottomMarginPct": 18,
"maxCharsPerLine": 22
}
The bottom margin matters more than people expect. Reels place the caption, the profile name and the action buttons along the bottom, so keeping text at roughly 18 percent above the bottom edge keeps your words clear of the interface. Set that once, reuse it on every reel.
Transcribe Your Reel and Get Accurate Word-Level Timings
Word-level timings are the difference between captions that feel glued to the audio and captions that feel approximate. Segment-level output gives you a phrase with one start time and one end time, which forces you to guess where each word sits. Word-level output tells you exactly.
Ask for the verbose JSON output, then flatten the word list into a simple array. You want a file that looks like this:
[
{"w": "most", "s": 0.42, "e": 0.71},
{"w": "people", "s": 0.71, "e": 1.14},
{"w": "scroll", "s": 1.14, "e": 1.62}
]
Then group the words into caption cards. A card is usually three to five words, and the split rule is simple: break at punctuation, break when the character count passes your limit, and never break in the middle of a two-word phrase people say together. Keep the group’s start time as the first word’s start and the end time as the last word’s end.
Two accuracy notes worth acting on. First, spell out numbers and brand names in a small replacement dictionary before rendering, because transcription models guess at proper nouns. Second, if your clip has music under the voice, run a vocal separation pass first. The transcript gets cleaner, and the timings get less jittery.
Style and Burn In Captions That Match Your Brand
Caption style is a brand asset. Pick a font you can license, a highlight color that appears elsewhere in your content, and one animation rule. Then never change it without a reason.
A practical default set:
- All caps, extra bold weight, one or two lines maximum on screen at a time.
- White text with a dark outline, plus a colored highlight on the single most important word in each card.
- A pop-in of about 80 to 120 milliseconds on card entry, no exit animation. Exit animations cost you legibility time.
- Consistent position. If your captions jump around the frame, viewers stop reading them.
For FFmpeg, an ASS subtitle file is the easiest path because it carries styling per line and per word. Generate the ASS file from your transcript and brand config, then burn it in as a single pass. In Remotion, you read the same transcript array and render a caption component frame by frame, which gives you smooth per-word highlighting without fighting subtitle syntax.
Whichever route you take, check the result on an actual phone at arm’s length. On a desktop monitor everything looks fine. On a phone in daylight, thin fonts disappear.
Export Reels-Ready Captions in the Right Aspect Ratio and Format
Reels are vertical. Render at 1080 by 1920, keep the frame rate consistent with your source, and use a high quality H.264 encode so the platform does not add another compression pass on top of a weak one.
A few export rules that save re-uploads:
- Export H.264 in an MP4 container. It is the format every network accepts.
- Keep the audio at 44.1 kHz or 48 kHz stereo. Mono files sometimes get rejected or sound thin after processing.
- Leave headroom above and below your captions. Cropping later breaks burned-in text permanently.
- Export one master file. Do not export a TikTok version and an Instagram version of the same reel unless the crops genuinely differ.
If you also post horizontal video to YouTube, keep a separate long-form master. Vertical clips become Shorts automatically, and custom thumbnails apply to long-form uploads rather than Shorts. The details on comments and pinning are in the YouTube auto pin comment guide.
Common Pitfalls: Sync Drift, Emoji, Fonts and Platform Limits
Sync drift. Drift usually comes from variable frame rate source files, not from the transcript. Normalize the clip to a constant frame rate before transcription so timings and video frames share one clock.
Emoji. Emoji render differently across fonts and operating systems, and some fonts silently drop them. If a caption card needs an emoji, render it as an image overlay instead of text, or restrict yourself to emoji your font definitely ships with.
Fonts. A font that looks right on your machine may not embed cleanly in the subtitle render. Convert text to outlines when the layout is final, or pick a font family with a permissive license and keep the file next to the project.
Platform limits. Instagram captions have no clickable links, so a link belongs in your first comment. Threads posts cap at 500 characters with a 250 post limit per 24 hours. X caps standard accounts at roughly 2 minutes 20 seconds of video. Knowing these before you export saves a re-cut.
It also helps to know what the platform itself will and will not attach. Trending audio cannot be attached through the API, so bake sound into the file. Trial Reels behave differently from standard Reels, which the Instagram trial reels API notes explain, and the Facebook Reels API guide covers the feed side of the same workflow.
Automate Captions Across Every Reel With Creator OS
Once you have a render script that produces a captioned file, the remaining work is publishing, captioning and reading results without opening seven dashboards. Creator OS is one place to publish, schedule, reply and read analytics on Instagram, TikTok, YouTube, X, LinkedIn, Facebook and Threads, plus Skool communities, a WordPress blog and paid ads as an add-on. It works from the web app, the iOS app, a REST API, a CLI and a hosted MCP server for Claude, ChatGPT, Claude Code, Cursor, Codex, Windsurf and VS Code.
The CLI is the natural fit for a caption pipeline because it is just another step in your script. Upload the rendered reel, then create the post with your caption, hashtags and first comment already in place.
creatoros media:upload out/reel-01-captioned.mp4
creatoros validate:post-length --text "Three fixes for jumpy captions"
creatoros posts:create --text "Three fixes for jumpy captions" \
--platforms instagram,tiktok,youtube \
--media med_123 \
--hashtags captions,reels,editing
Two details make this reliable. Uploading media once returns a med_ id you can reuse on every platform, so you are not re-uploading the same file four times. And if you name a network in a post that is not connected, the API returns it in missing_platforms while the rest still go out. Your batch does not fail because one account is stale.
For an agent-driven version, the hosted MCP server at https://mcp.creatoros.ca/mcp exposes tools including create_post, upload_media_from_url, get_best_time_to_post, check_caption_length and get_post_analytics. Read-only connections only see read tools, and destructive tools are marked so the AI app asks before it acts. Setting it up takes one connector entry, described at the MCP docs.
There is also a ready-made harness. Social Agents is an open-source agent that runs your socials on Creator OS: it interviews you about your brand, then posts, replies to comments and DMs, runs automations and reports from a local dashboard. You bring your own model, either a logged-in Claude Code session, an ANTHROPIC_API_KEY, or any Anthropic-compatible API. The repository is at github.com/kevinbadi/social-agents, and the background is in the Social Agents open source post.
If you would rather do the captioning inside Claude Code, npx @creatoros/cli@latest init installs 14 skills, including post-shortform, post-everywhere and analytics. creatoros sync updates them later without overwriting files you have edited, since updates land as .new files instead.
Watch the walkthroughs on the KevBuildsApps YouTube channel, and if you want to see the local render pipeline running against a real reel, the launch video is at youtu.be/-QBJH_PK3pY.
Get Started With Creator OS
The Creator plan is $19.99/month or $59.99/year and covers up to 8 connected accounts: 7 socials plus Skool. API keys, the MCP server, the CLI and the agent skills are included in every plan, so your caption script can publish the moment the render finishes.
Sign up at creatoros.ca/sign-up. For endpoint details, request shapes and the full tool list, start at the Creator OS docs.