Video Understand

calesthio/OpenMontage/.claude/skills/video-understand

作者 calesthio9327439db69021ab4b0e2776729bf3b58fdb5a87無授權條款65K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫5 天前更新

Understand video content locally using ffmpeg frame extraction and Whisper transcription. No API keys needed. Use when: (1) Understanding what a video contains, (2) Transcribing video audio locally, (3) Extracting key frames for visual analysis, (4) Getting video content without API keys.

包含腳本Design & Creative
AI 產生的概覽

使用 ffmpeg 擷取影片畫面並以 Whisper 在本機轉寫音訊,輸出 JSON。

功能
執行一支 Python 指令碼,使用 ffmpeg 以場景、關鍵畫面或固定間隔模式從影片擷取畫面,並可選擇以 Whisper 轉寫音訊。它會輸出包含影片中繼資料、帶時間戳的畫面路徑與轉寫片段的 JSON,也可將該 JSON 寫入檔案。之後可由代理檢視畫面影像進行視覺理解。
適用情境
適合用來了解影片內容、離線轉寫影片音訊,或擷取關鍵畫面進行視覺分析,且不需要 API 金鑰。
執行需求
需要 ffmpeg 與 ffprobe;openai-whisper 為選用,僅在轉寫時需要。此技能附帶可執行的 Python 指令碼,完全在本機執行,不需要 API 金鑰或網路存取。

video-understand

Understand video content locally using ffmpeg for frame extraction and Whisper for transcription. Fully offline, no API keys required.

Prerequisites

  • ffmpeg + ffprobe (required): brew install ffmpeg
  • openai-whisper (optional, for transcription): pip install openai-whisper

Commands

bash
# Scene detection + transcribe (default)python3 skills/video-understand/scripts/understand_video.py video.mp4
# Keyframe extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m keyframe
# Regular interval extractionpython3 skills/video-understand/scripts/understand_video.py video.mp4 -m interval
# Limit frames extractedpython3 skills/video-understand/scripts/understand_video.py video.mp4 --max-frames 10
# Use a larger Whisper modelpython3 skills/video-understand/scripts/understand_video.py video.mp4 --whisper-model small
# Frames only, skip transcriptionpython3 skills/video-understand/scripts/understand_video.py video.mp4 --no-transcribe
# Quiet mode (JSON only, no progress)python3 skills/video-understand/scripts/understand_video.py video.mp4 -q
# Output to filepython3 skills/video-understand/scripts/understand_video.py video.mp4 -o result.json

CLI Options

FlagDescription
videoInput video file (positional, required)
-m, --modeExtraction mode: scene (default), keyframe, interval
--max-framesMaximum frames to keep (default: 20)
--whisper-modelWhisper model size: tiny, base, small, medium, large (default: base)
--no-transcribeSkip audio transcription, extract frames only
-o, --outputWrite result JSON to file instead of stdout
-q, --quietSuppress progress messages, output only JSON

Extraction Modes

ModeHow it worksBest for
sceneDetects scene changes via ffmpeg select='gt(scene,0.3)'Most videos, varied content
keyframeExtracts I-frames (codec keyframes)Encoded video with natural keyframe placement
intervalEvenly spaced frames based on duration and max-framesFixed sampling, predictable output

If scene mode detects no scene changes, it automatically falls back to interval mode.

Output

The script outputs JSON to stdout (or file with -o). See references/output-format.md for the full schema.

json
{  "video": "video.mp4",  "duration": 18.076,  "resolution": {"width": 1224, "height": 1080},  "mode": "scene",  "frames": [    {"path": "/abs/path/frame_0001.jpg", "timestamp": 0.0, "timestamp_formatted": "00:00"}  ],  "frame_count": 12,  "transcript": [    {"start": 0.0, "end": 2.5, "text": "Hello and welcome..."}  ],  "text": "Full transcript...",  "note": "Use the Read tool to view frame images for visual understanding."}

Use the Read tool on frame image paths to visually inspect extracted frames.

References

  • references/output-format.md -- Full JSON output schema documentation

來源與署名

來源:calesthio/OpenMontage位於.claude/skills/video-understand提交9327439

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架