Boson Video

io.github.linboxinv0.2.0更新于 Oct 8, 2026

Let your AI read any video: timed transcript, full-resolution frames, screen text, claim checks.

概览

AI 生成的概览

让助手读懂视频:带时间点的字幕、全分辨率画面、屏幕文字与说法核查。

功能
根据 YouTube 链接或本地文件为视频生成带时间点的文档,涵盖场景、章节、字幕、屏幕文字和摘要。助手可以搜索视频、查看全分辨率画面,并给出带时间点引用的回答。它还会把摘要句子与实际内容比对,标记为已确认、存疑或错误。
适用场景
适合长演讲、课程和财经视频,尤其是用户只略懂其语言的场景,需要基于具体时间点而非完整观看得到答案时。若助手需要引用或核实视频内容,值得添加。
运行要求
以 Python 包形式通过 uvx 在本地运行;场景、字幕和搜索无需账号或密钥。可选的 INCEPTION_API_KEY 和 TYPESAFE_API_KEY 可启用更完整的摘要、翻译和说法核查。语音模型在首次使用时下载。仅支持桌面端。
安装前请注意
可选的 TYPESAFE_API_KEY 属于机密,应以环境变量方式提供,不要外泄。处理时会把视频内容交给本地模型;设置可选密钥后,还会发送给相应的第三方服务。README 指出 YouTube 条款不允许产品化的自动访问,因此使用应限于用户自己的机器。

安装

在 SourceWeft 中

  1. 打开 控制台中的 Boson Video,将其添加到工作区。
  2. 为需要使用其工具的对话启用该服务。

Desktop only,通过 STDIO。 STDIO 服务会启动本地进程,因此需要 SourceWeft 桌面宿主。

其他 MCP 客户端

参照 仓库 中的启动说明。

README

Boson-Video

Get the point of any video without watching all of it. Paste a YouTube link or drop a video file. Every line of the summary, the transcript and the screen text carries the second it came from, and questions are answered from the video and checked against it.

[A video read in Boson-Video: the player and the ribbon on the left, the terms explained on the right]

  • The ribbon: the whole video on one strip (scenes, chapters, most replayed). Hover to see any second's frame and words.
  • Summary: sections with their frames; every sentence timed and checked (✓ ? ✗).
  • Transcript: the original with English beneath, technical terms explained, on-screen text read in.
  • Ask: answers that cite their moments and show the frames, with background kept apart.
  • Your own AI: the same document as an MCP plugin for Claude, Cursor, Codex and more.

Built for long talks, lectures and finance videos in a language you half know (Chinese and English today).

In your own AI

One line, nothing else to install (uv runs it):

bash
claude mcp add boson-video -- uvx boson-video mcp

Claude Desktop (claude_desktop_config.json) or Cursor (.cursor/mcp.json):

json
{ "mcpServers": { "boson-video": { "command": "uvx", "args": ["boson-video", "mcp"] } } }

Codex (~/.codex/config.toml):

toml
[mcp_servers.boson-video]command = "uvx"args = ["boson-video", "mcp"]

Then ask your AI about any YouTube link. It opens the video, reads the transcript, looks at the frames that matter at full resolution, searches, and checks its claims, citing every moment. No keys needed. Everything runs on your computer. Details: docs/plugin.md.

Run the page

bash
uvx boson-video web             # http://127.0.0.1:8770; the first run prints your invite code

Keys make it fuller: INCEPTION_API_KEY (Mercury writes the summary, the English and the terms) and TYPESAFE_API_KEY (Jev checks every sentence and answers questions). Without them you still get the scenes, the transcript and search.

Tested videos

Measured, one run each unless a range is shown. Mac: M5 MacBook with Apple's transcriber. Windows: a 12-thread laptop with SenseVoice or Parakeet on the CPU. Full tables: docs/measured.md.

VideoLengthScene mapWordsTranscript errors¹Screen textSummary ✓
Money or Life 美股频道, Meta and AI (zh, talking head)24:531.3 s13.8 s (Mac)no human captions—13 / 13
Money or Life 美股频道, AI drug discovery (zh, slides)28:241.1 s18.5 s (Mac)not scored36 moments read15 / 16
程序员老王, LLM abliteration (zh, animated diagrams)11:35—10.0 s (Windows)not scored88 moments; subtitles apart at 8521–22 / 22–23 (several runs)
陳永儀, TEDxTaipei (zh, talk)14:29—7.1 s (Mac)5.3% chars (Mac), 4.8% (SenseVoice)—14 / 14
Ken Robinson, TED (en, talk)20:06—9.9 s (Mac)10.0% words (Mac), 10.5% (Parakeet)—22 / 25
Sean's AI Stories, agent observability (en, screen recording)20:481.5 s13.2 s (Mac)not scorednone: YouTube refused full resolution23 / 26
Andrej Karpathy, Intro to LLMs (en, slides)59:481.6 s30.5 s (Mac)only automatic captions19 of 20 slides right22 / 23
Rick Astley, Never Gonna Give You Up (fast cuts)3:330.7 s————

¹ Against human captions (scripts/accuracy.py); about half the English errors are filler words the captions leave out. Summary ✓: sentences confirmed against what was said (Jev and code).

Command line

bash
uvx boson-video <link | id | file>        # build a video: scenes, words, screen text, summaryuvx boson-video ask <video> "question"    # the moment that answers ituvx boson-video models                    # fetch the local speech models now instead of on first use

From a clone: uv sync, then uv run boson-video …; uv run pytest runs the tests (offline).

More

MIT licensed. YouTube's terms don't allow automated access for products, so everything that touches YouTube runs on your own computer.

来源:README.md,提交 cc587a8

工具

0
工具元数据尚未被收录。

版本历史

1
  1. v0.2.0最新Oct 8, 2026