Video Understand

calesthio/OpenMontage/.agents/skills/video-understand

by calesthio9327439db69021ab4b0e2776729bf3b58fdb5a87No license65K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 5 days ago

Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

Includes scriptsDesign & Creative
AI-generated overview

Extracts video frames with ffmpeg and transcribes audio locally with Whisper, returning JSON.

What it does
Runs a Python script that uses ffmpeg to extract frames from a video in scene, keyframe, or interval mode, and optionally transcribes the audio with Whisper. It outputs JSON containing video metadata, frame paths with timestamps, and transcript segments, or writes that JSON to a file. Frame images can then be inspected visually by the agent.
When to use it
Use it to understand what a video contains, transcribe video audio offline, or pull key frames for visual analysis without API keys.
Requirements
Requires ffmpeg and ffprobe; openai-whisper is optional and only needed for transcription. It ships an executable Python script and runs locally with no API keys or network access.

video-understand

Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.

Prerequisites

  • ffmpeg + ffprobe (required): brew install ffmpeg
  • openai-whisper (optional, for transcription): pip install openai-whisper

Commands

bash
# Scene detection + transcribe (default)python3 skills/video-understand/scripts/understand_video.py video.mp4
# Keyframe extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe
# Regular interval extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval
# Limit frames extractedpython3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10
# Use a larger Whisper modelpython3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small
# Frames only, skip transcriptionpython3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe
# Quiet mode (JSON only, no progress)python3 skills/video-understand/scripts/understand_video.py video.mp4 -q
# Output to filepython3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json

CLI Options

FlagDescription
videoInput video file (positional, required)
-m, --modeExtraction mode: scene (default), keyframe, interval
--max-framesMaximum frames to keep (default: 20)
--whisper-modelWhisper model size: tiny, base, small, medium, large (default: base)
--no-transcribeSkip audio transcription, extract frames only
-o, --outputWrite result JSON to file instead of stdout
-q, --quietSuppress progress messages, output only JSON

Extraction Modes

ModeHow it worksBest for
sceneDetects scene changes via ffmpeg select='gt(scene,0.3)'Most videos, varied content
keyframeExtracts I-frames (codec keyframes)Encoded video with natural keyframe placement
intervalEvenly spaced frames based on duration and max-framesFixed sampling, predictable output

If scene mode detects no scene changes, it automatically falls back to interval mode.

Output

The script outputs JSON to stdout (or file with -o). See references/output-format.md for the full schema.

json
{  "video": "video.mp4",  "duration": 18.076,  "resolution": {"width": 1224, "height": 1080},  "mode": "scene",  "frames": [    {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"}  ],  "frame_count": 12,  "transcript": [    {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."}  ],  "text": "Full transcript...",  "note": "Use the Read tool to view frame images for visual understanding."}

Use the Read tool on frame image paths to visually inspect extracted frames.

References

  • references/output-format.md -- Full JSON output schema documentation

Source and attribution

Source:calesthio/OpenMontagein.agents/skills/video-understandat commit9327439

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal