Video Understand

calesthio/OpenMontage/.agents/skills/video-understand

作者 calesthio9327439db69021ab4b0e2776729bf3b58fdb5a87无许可证65K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库5天前更新

Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

包含脚本Design & Creative
AI 生成的概览

使用 ffmpeg 提取视频帧并用 Whisper 在本地转写音频,输出 JSON。

功能
运行一个 Python 脚本,使用 ffmpeg 以场景、关键帧或固定间隔模式从视频中提取帧,并可选地用 Whisper 转写音频。它输出包含视频元数据、带时间戳的帧路径和转写片段的 JSON,也可将该 JSON 写入文件。随后可由智能体查看帧图像进行视觉理解。
适用场景
适用于了解视频内容、离线转写视频音频,或提取关键帧进行视觉分析,且无需 API 密钥。
运行要求
需要 ffmpeg 和 ffprobe;openai-whisper 为可选,仅在转写时需要。该技能附带可执行的 Python 脚本,完全在本地运行,无需 API 密钥或网络访问。

video-understand

Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.

Prerequisites

  • ffmpeg + ffprobe (required): brew install ffmpeg
  • openai-whisper (optional, for transcription): pip install openai-whisper

Commands

bash
# Scene detection + transcribe (default)python3 skills/video-understand/scripts/understand_video.py video.mp4
# Keyframe extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe
# Regular interval extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval
# Limit frames extractedpython3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10
# Use a larger Whisper modelpython3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small
# Frames only, skip transcriptionpython3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe
# Quiet mode (JSON only, no progress)python3 skills/video-understand/scripts/understand_video.py video.mp4 -q
# Output to filepython3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json

CLI Options

FlagDescription
videoInput video file (positional, required)
-m, --modeExtraction mode: scene (default), keyframe, interval
--max-framesMaximum frames to keep (default: 20)
--whisper-modelWhisper model size: tiny, base, small, medium, large (default: base)
--no-transcribeSkip audio transcription, extract frames only
-o, --outputWrite result JSON to file instead of stdout
-q, --quietSuppress progress messages, output only JSON

Extraction Modes

ModeHow it worksBest for
sceneDetects scene changes via ffmpeg select='gt(scene,0.3)'Most videos, varied content
keyframeExtracts I-frames (codec keyframes)Encoded video with natural keyframe placement
intervalEvenly spaced frames based on duration and max-framesFixed sampling, predictable output

If scene mode detects no scene changes, it automatically falls back to interval mode.

Output

The script outputs JSON to stdout (or file with -o). See references/output-format.md for the full schema.

json
{  "video": "video.mp4",  "duration": 18.076,  "resolution": {"width": 1224, "height": 1080},  "mode": "scene",  "frames": [    {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"}  ],  "frame_count": 12,  "transcript": [    {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."}  ],  "text": "Full transcript...",  "note": "Use the Read tool to view frame images for visual understanding."}

Use the Read tool on frame image paths to visually inspect extracted frames.

References

  • references/output-format.md -- Full JSON output schema documentation

来源与署名

来源:calesthio/OpenMontage位于.agents/skills/video-understand提交9327439

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架