Boson Video

io.github.linboxinv0.2.0Updated Oct 8, 2026

Let your AI read any video: timed transcript, full-resolution frames, screen text, claim checks.

Overview

AI-generated overview

Lets an assistant read a video: timed transcript, full-resolution frames, on-screen text, and claim checks.

What it does
Builds a timed document for a video from a YouTube link or a local file, covering scenes, chapters, transcript, on-screen text and a summary. The assistant can search the video, look at frames at full resolution, and answer questions with citations to the exact moment. It also checks summary sentences against what was actually said, marking them confirmed, uncertain or wrong.
When to use it
Useful for long talks, lectures and finance videos, especially in a language the user only partly knows, when you want answers grounded in specific moments rather than a whole viewing. Worth adding if the assistant needs to quote or verify what a video says.
Requirements
Runs locally as a Python package launched with uvx; no account or key is required for scenes, transcript and search. Optional keys INCEPTION_API_KEY and TYPESAFE_API_KEY unlock fuller summaries, translation and claim checking. Speech models are downloaded on first use. Desktop only.
Before you install
The optional TYPESAFE_API_KEY is a secret and should be supplied as an environment variable, not shared. Processing sends video content to local models and, when the optional keys are set, to the corresponding third-party services. The README notes YouTube's terms do not allow automated access for products, so use is intended to stay on the user's own machine.

Installation

In SourceWeft

  1. Open Boson Video in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

Boson-Video

Get the point of any video without watching all of it. Paste a YouTube link or drop a video file. Every line of the summary, the transcript and the screen text carries the second it came from, and questions are answered from the video and checked against it.

[A video read in Boson-Video: the player and the ribbon on the left, the terms explained on the right]

  • The ribbon: the whole video on one strip (scenes, chapters, most replayed). Hover to see any second's frame and words.
  • Summary: sections with their frames; every sentence timed and checked (✓ ? ✗).
  • Transcript: the original with English beneath, technical terms explained, on-screen text read in.
  • Ask: answers that cite their moments and show the frames, with background kept apart.
  • Your own AI: the same document as an MCP plugin for Claude, Cursor, Codex and more.

Built for long talks, lectures and finance videos in a language you half know (Chinese and English today).

In your own AI

One line, nothing else to install (uv runs it):

bash
claude mcp add boson-video -- uvx boson-video mcp

Claude Desktop (claude_desktop_config.json) or Cursor (.cursor/mcp.json):

json
{ "mcpServers": { "boson-video": { "command": "uvx", "args": ["boson-video", "mcp"] } } }

Codex (~/.codex/config.toml):

toml
[mcp_servers.boson-video]command = "uvx"args = ["boson-video", "mcp"]

Then ask your AI about any YouTube link. It opens the video, reads the transcript, looks at the frames that matter at full resolution, searches, and checks its claims, citing every moment. No keys needed. Everything runs on your computer. Details: docs/plugin.md.

Run the page

bash
uvx boson-video web             # http://127.0.0.1:8770; the first run prints your invite code

Keys make it fuller: INCEPTION_API_KEY (Mercury writes the summary, the English and the terms) and TYPESAFE_API_KEY (Jev checks every sentence and answers questions). Without them you still get the scenes, the transcript and search.

Tested videos

Measured, one run each unless a range is shown. Mac: M5 MacBook with Apple's transcriber. Windows: a 12-thread laptop with SenseVoice or Parakeet on the CPU. Full tables: docs/measured.md.

VideoLengthScene mapWordsTranscript errors¹Screen textSummary ✓
Money or Life 美股频道, Meta and AI (zh, talking head)24:531.3 s13.8 s (Mac)no human captions—13 / 13
Money or Life 美股频道, AI drug discovery (zh, slides)28:241.1 s18.5 s (Mac)not scored36 moments read15 / 16
程序员老王, LLM abliteration (zh, animated diagrams)11:35—10.0 s (Windows)not scored88 moments; subtitles apart at 8521–22 / 22–23 (several runs)
陳永儀, TEDxTaipei (zh, talk)14:29—7.1 s (Mac)5.3% chars (Mac), 4.8% (SenseVoice)—14 / 14
Ken Robinson, TED (en, talk)20:06—9.9 s (Mac)10.0% words (Mac), 10.5% (Parakeet)—22 / 25
Sean's AI Stories, agent observability (en, screen recording)20:481.5 s13.2 s (Mac)not scorednone: YouTube refused full resolution23 / 26
Andrej Karpathy, Intro to LLMs (en, slides)59:481.6 s30.5 s (Mac)only automatic captions19 of 20 slides right22 / 23
Rick Astley, Never Gonna Give You Up (fast cuts)3:330.7 s————

¹ Against human captions (scripts/accuracy.py); about half the English errors are filler words the captions leave out. Summary ✓: sentences confirmed against what was said (Jev and code).

Command line

bash
uvx boson-video <link | id | file>        # build a video: scenes, words, screen text, summaryuvx boson-video ask <video> "question"    # the moment that answers ituvx boson-video models                    # fetch the local speech models now instead of on first use

From a clone: uv sync, then uv run boson-video …; uv run pytest runs the tests (offline).

More

MIT licensed. YouTube's terms don't allow automated access for products, so everything that touches YouTube runs on your own computer.

Source: README.md at commit cc587a8

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v0.2.0LatestOct 8, 2026