
Justcaptions
io.github.SpaceSheepBoyv1.3.1Updated Oct 4, 2026
Audio transcription, caption correction, translation and 15 animated video caption styles.
Overview
Transcribes video or audio, corrects and translates captions, and burns animated caption styles into videos locally or via a cloud API.
- What it does
- Just Captions transcribes speech and produces SRT, VTT and JSON caption files, then burns animated TikTok/Reels-style captions into video. It offers 15 animated presets (Emoji, Word Highlight, Neon, Typewriter and others) with configurable colors, fonts, position, size and safe-area margins. Local tools include check_environment, list_styles, preview_style, caption_video, get_job, get_usage and job recovery; the remote endpoint adds transcribe_audio, correct_captions, translate_captions, pick_emojis, get_pricing and get_usage. It handles one video or a folder of up to 100 files.
- When to use it
- Use it when you want an assistant to caption or subtitle videos, generate SRT/VTT/JSON files, translate or correct existing captions, or apply social-media caption styles in batch. It suits creators and developers automating video captioning from an MCP client or the command line.
- Requirements
- The packaged local MCP/CLI needs Python 3.10+, uv and ffmpeg; the plain skill needs Python 3.9+ and Pillow. The remote endpoint at uses Streamable HTTP with an Authorization header containing a Bearer API key, supplied via JUSTCAPTIONS_API_KEY or a saved key file. Offline transcription needs faster-whisper, which downloads a model on first use. Cloud transcription, correction, translation and emoji picking require an API key.
Installation
In SourceWeft
- Open Justcaptions in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"justcaptions": {
"type": "http",
"url": "https://api.justcaptions.com/mcp"
}
}
}README
Just Captions for agents
Caption videos from Codex, Claude Code, any MCP client, or your terminal. This skill transcribes the speech, makes SRT/VTT/JSON files, and burns TikTok/Reels-style captions into the video. It does one video or a whole folder at a time.
15 animated portable presets follow the Just Captions app style families. Typography and sampled animations can differ from native iOS. Rendering runs on your machine with Pillow and ffmpeg.
Choose a style and copy a task · API docs · OpenAPI · Style catalog
MCP quickstart
Python 3.10+, uv and ffmpeg are required for the packaged MCP/CLI. The plain skill still works with Python 3.9+ and Pillow.
Codex — add to ~/.codex/config.toml:
Claude Code — merge into your project's .mcp.json:
Set JUSTCAPTIONS_API_KEY, or save a key with the CLI's --signup YOUR_EMAIL. The local MCP also reads ~/.config/justcaptions/api_key. Never put keys in prompts or commits. For recognition without a key, add "--with", "faster-whisper" before "--from" in the uvx arguments. The offline model downloads on first use.
Ask: “Caption ~/Movies/intro.mp4 with Word Highlight for TikTok. Save MP4, SRT, VTT and JSON in ./captioned/, and preserve the original.”
caption_video takes input_path, output_dir, style_id, size, position, safe_area, length, overrides, captions_path, engine, language, translate_to, correct, glossary, burn, and overwrite. See MCP discovery for typed schemas. Up to 100 files per folder job, two concurrent jobs. Jobs persist locally across process restarts; use list_jobs and resume_job to recover. Existing outputs are protected unless explicitly overwritten; folder jobs use per-source output directories.
Remote MCP
https://api.justcaptions.com/mcp uses Streamable HTTP and your API key. It handles audio and caption text. It does not access local files, preview styles or render MP4.
For Codex:
For Claude Code, a project .mcp.json entry:
Tools: list_styles, get_pricing, get_usage, transcribe_audio, correct_captions, translate_captions, pick_emojis. Cloud operations use the existing API pricing and limits. Supply a unique request_id and reuse it only for identical retries. Request hashes and completed replies have a 24-hour replay window, then are removed on a subsequent request or daily cleanup (within 48 hours). The replay store does not retain audio.
Styles and customization
All 15 presets animate in video exports and the website previews, using spoken-word highlights/reveals, pop entrances, fades, typewriting or a pulsing glow.
The authoritative versioned catalog is skills/justcaptions/assets/styles.json: Emoji, Mega, Reveal, Neon, White box, Yellow box, Gray box, Yellow outline, Word Highlight, Highlight Box, Impact, Pop In, Typewriter, Cinema and Editorial. Historical karaoke, black-box and white-outline CLI names remain supported. App IDs such as wordHighlight resolve to their portable preset IDs.
The local MCP accepts overrides, for example:
The CLI takes the same object in --style-config FILE.json. Unknown keys and invalid values are rejected. Colors, backgrounds, outlines, font families, size, word count and letter case are configurable. Use --safe-area tiktok|reels|shorts|none; conservative margins keep captions horizontally centered. Platform UI layouts can vary.
Recognition word timestamps are retained when text is unchanged. Corrected/translated text and imported subtitle-only files use estimated word timing, not forced alignment. Serif/regular families use available system fonts, with a bundled Geist fallback.
Install
You need ffmpeg and Python 3.9+ with Pillow:
As a Claude Code plugin
As a plain skill
Then ask Claude something like "caption interview.mp4 with the emoji style" or "burn Spanish subtitles into everything in ./clips".
Use it without Claude
To change the wording before you burn: run the script without --burn, edit NAME.json or NAME.srt, then run it again with --captions NAME.json --burn. The JSON file keeps per-word timing for karaoke and emoji.
Just Captions API
Without a key, everything runs offline with faster-whisper. With a key you get:
- cloud transcription (better on accents, noisy audio and many languages)
--correctand--translate- AI emoji picks for the Emoji style (the offline fallback is a keyword table)
Getting a key is instant and needs no card:
The key is saved to ~/.config/justcaptions/api_key (readable only by you). If JUSTCAPTIONS_API_KEY is set, it is used instead.
Pricing
- Free plan (no card): stops when the free allowance runs out. When faster-whisper is installed, transcription carries on locally.
- Pay as you go: add a card at https://justcaptions.com/api/account/ to keep going past the free allowance. Stripe invoices you monthly, and the default spend cap is $100 a month.
- Failed requests are not billed.
The API takes audio only. The script pulls a small mono track out of your video, so the video itself is never uploaded, and burning always happens on your machine. API reference: https://justcaptions.com/api/
Want this on your iPhone?
Just Captions on the App Store provides native caption styling on your phone, with on-device transcription, editing and batch export.
Development
Install in an isolated environment with uv pip install -e .. The worker MCP adapter is open source in remote/mcp.mjs; the host injects its existing authenticated API handler. tools/sync-agent-assets.py --site SITE_DIR --backend BACKEND_DIR publishes the canonical catalog, OpenAPI and adapter to the site and Worker source. Generated copies should not be edited independently.
How burning works: each caption state (one per word for karaoke and emoji) is drawn as a full-frame transparent PNG. The concat demuxer plays the PNGs back as a single image stream with exact durations, and one overlay filter puts it over the video. That keeps the ffmpeg command the same size no matter how many captions there are.
License
The code is MIT. The bundled Geist font is SIL OFL 1.1 (skills/justcaptions/assets/fonts/OFL.txt).
Complete first export
After installing Python 3.10+, ffmpeg and the package, run justcaptions --demo --out-dir ./demo --style word-highlight. The bundled human-narration sample includes word timestamps and needs no API key or model download. In MCP use run_demo, then get_job. Narration is attributed in demo credits.
Durable jobs and account controls
list_jobs, get_job, cancel_job and resume_job operate on saved jobs in ~/.local/share/justcaptions (override with JUSTCAPTIONS_STATE_DIR). Resume skips completed files and reuses successful cloud responses. Source files must remain unchanged. An in-flight cloud request can finish and be charged after cancellation. Uncertain cloud requests older than the server replay window require manual usage review rather than silently repeating paid work.
Use estimate_video before recognition, supplying known text_chars for edits. Estimates exclude unknown future text and do not reserve credit. The API reserves capacity atomically for concurrent requests and all dedicated keys share one account's allowance. Manage monthly spend limits, scoped keys, expiry and revocation at https://justcaptions.com/api/account/. Existing charges and running holds are not removed by lowering a limit.
Setup/export statistics are optional. JUSTCAPTIONS_METRICS_ID enables the counters included in copied configurations only when a website visitor opts in. No file paths, captions or API keys are sent by these counters. To disable an installed client's reporting, remove that variable and restart the MCP server.
Distribution
Install a standard wheel from the latest GitHub release:
GitHub Actions builds and checks wheel/sdist artifacts, attaches them to releases and publishes server.json to the official MCP Registry using GitHub OIDC. PyPI publishing uses a trusted publisher when the owner enables it; the website uses the release wheel until that registration is completed.
PyPI pending trusted publisher configuration: project justcaptions-agent; GitHub owner SpaceSheepBoy; repository justcaptions-skill; workflow release.yml; environment pypi. After registration, enable repository variable PYPI_PUBLISH_ENABLED=true and rerun the release workflow.
Source: README.md at commit 1ea77d2
Tools
0Version history
1- v1.3.1LatestOct 4, 2026


