Boson Video

io.github.linboxinv0.2.0更新於 Oct 8, 2026

Let your AI read any video: timed transcript, full-resolution frames, screen text, claim checks.

概覽

AI 產生的概覽

讓助理讀懂影片:帶時間點的逐字稿、全解析度畫面、畫面文字與說法查核。

功能
依 YouTube 連結或本機檔案為影片建立帶時間點的文件,涵蓋場景、章節、逐字稿、畫面文字與摘要。助理可以搜尋影片、檢視全解析度畫面,並給出附時間點引用的回答。它也會把摘要句子與實際內容比對,標示為已確認、存疑或錯誤。
適用情境
適合長演講、課程與財經影片,尤其是使用者只略懂該語言的場合,需要依具體時間點而非完整觀看取得答案時。若助理需要引用或查核影片內容,值得加入。
執行需求
以 Python 套件形式透過 uvx 在本機執行;場景、逐字稿與搜尋不需要帳號或金鑰。選用的 INCEPTION_API_KEY 與 TYPESAFE_API_KEY 可啟用更完整的摘要、翻譯與說法查核。語音模型於首次使用時下載。僅支援桌面端。
安裝前請注意
選用的 TYPESAFE_API_KEY 屬於機密,應以環境變數提供,不要外流。處理時會把影片內容交給本機模型;設定選用金鑰後,也會傳送給對應的第三方服務。README 指出 YouTube 條款不允許產品化的自動存取,因此使用應限於使用者自己的機器。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 Boson Video,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

Boson-Video

Get the point of any video without watching all of it. Paste a YouTube link or drop a video file. Every line of the summary, the transcript and the screen text carries the second it came from, and questions are answered from the video and checked against it.

[A video read in Boson-Video: the player and the ribbon on the left, the terms explained on the right]

  • The ribbon: the whole video on one strip (scenes, chapters, most replayed). Hover to see any second's frame and words.
  • Summary: sections with their frames; every sentence timed and checked (✓ ? ✗).
  • Transcript: the original with English beneath, technical terms explained, on-screen text read in.
  • Ask: answers that cite their moments and show the frames, with background kept apart.
  • Your own AI: the same document as an MCP plugin for Claude, Cursor, Codex and more.

Built for long talks, lectures and finance videos in a language you half know (Chinese and English today).

In your own AI

One line, nothing else to install (uv runs it):

bash
claude mcp add boson-video -- uvx boson-video mcp

Claude Desktop (claude_desktop_config.json) or Cursor (.cursor/mcp.json):

json
{ "mcpServers": { "boson-video": { "command": "uvx", "args": ["boson-video", "mcp"] } } }

Codex (~/.codex/config.toml):

toml
[mcp_servers.boson-video]command = "uvx"args = ["boson-video", "mcp"]

Then ask your AI about any YouTube link. It opens the video, reads the transcript, looks at the frames that matter at full resolution, searches, and checks its claims, citing every moment. No keys needed. Everything runs on your computer. Details: docs/plugin.md.

Run the page

bash
uvx boson-video web             # http://127.0.0.1:8770; the first run prints your invite code

Keys make it fuller: INCEPTION_API_KEY (Mercury writes the summary, the English and the terms) and TYPESAFE_API_KEY (Jev checks every sentence and answers questions). Without them you still get the scenes, the transcript and search.

Tested videos

Measured, one run each unless a range is shown. Mac: M5 MacBook with Apple's transcriber. Windows: a 12-thread laptop with SenseVoice or Parakeet on the CPU. Full tables: docs/measured.md.

VideoLengthScene mapWordsTranscript errors¹Screen textSummary ✓
Money or Life 美股频道, Meta and AI (zh, talking head)24:531.3 s13.8 s (Mac)no human captions—13 / 13
Money or Life 美股频道, AI drug discovery (zh, slides)28:241.1 s18.5 s (Mac)not scored36 moments read15 / 16
程序员老王, LLM abliteration (zh, animated diagrams)11:35—10.0 s (Windows)not scored88 moments; subtitles apart at 8521–22 / 22–23 (several runs)
陳永儀, TEDxTaipei (zh, talk)14:29—7.1 s (Mac)5.3% chars (Mac), 4.8% (SenseVoice)—14 / 14
Ken Robinson, TED (en, talk)20:06—9.9 s (Mac)10.0% words (Mac), 10.5% (Parakeet)—22 / 25
Sean's AI Stories, agent observability (en, screen recording)20:481.5 s13.2 s (Mac)not scorednone: YouTube refused full resolution23 / 26
Andrej Karpathy, Intro to LLMs (en, slides)59:481.6 s30.5 s (Mac)only automatic captions19 of 20 slides right22 / 23
Rick Astley, Never Gonna Give You Up (fast cuts)3:330.7 s————

¹ Against human captions (scripts/accuracy.py); about half the English errors are filler words the captions leave out. Summary ✓: sentences confirmed against what was said (Jev and code).

Command line

bash
uvx boson-video <link | id | file>        # build a video: scenes, words, screen text, summaryuvx boson-video ask <video> "question"    # the moment that answers ituvx boson-video models                    # fetch the local speech models now instead of on first use

From a clone: uv sync, then uv run boson-video …; uv run pytest runs the tests (offline).

More

MIT licensed. YouTube's terms don't allow automated access for products, so everything that touches YouTube runs on your own computer.

來源:README.md,提交 cc587a8

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.2.0最新Oct 8, 2026