Scribiz
com.scribizv0.1.0Updated Oct 5, 2026
Transcripts, summaries, chapters and timestamped answers for video links, for AI agents.
Overview
Lets an assistant pull transcripts, summaries, chapters and timestamped answers from public video links.
- What it does
- Scribiz is a hosted MCP server that gives an assistant context for a public video link. Its five read-only tools return a title, length, summary, chapters and key moments with links (get_video_context), search a video for spoken or shown moments (search_video), answer a question with cited moments (ask_video), page through a transcript in TXT, Markdown, SRT, VTT or JSON (get_transcript), and check a video still being processed (get_job). Each result reports which layers ran and how many minutes it used.
- When to use it
- Use it when you want an assistant to summarize a video, find the exact moment a topic is mentioned, quote a passage, or answer questions about a video without reading the whole transcript. It suits chat and coding agents that work with public video links.
- Requirements
- A remote endpoint at over streamable HTTP; no account or key is needed for the free tier. An optional Scribiz API key, sent as an Authorization: Bearer header, unlocks Watch (on-screen reading) and uses plan minutes. Network access is required; the hosted server takes public links only and cannot read local files.
Installation
In SourceWeft
- Open Scribiz in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Web executable via Streamable HTTP. Remote servers run from the web runtime once configured in a workspace.
Other MCP clients
Add this to your client's mcpServers config.
{
"mcpServers": {
"mcp": {
"type": "http",
"url": "https://scribiz.com/mcp"
}
}
}README
Scribiz MCP server
Give your AI agent the context of a video: the transcript, a summary, chapters, and the exact moments that answer a question, each with a timestamp link.
This repo holds the setup files for the hosted Scribiz MCP server at https://scribiz.com/mcp. The server runs on scribiz.com. The Scribiz source code is not in this repo.
What the agent does with it
You paste a video link into a chat. The agent asks Scribiz for a short overview first. It searches for the part it needs, reads only that part, and answers with a link that opens the video at that moment. It does not read a whole transcript to answer one question.
Add it
The hosted server needs no account and no key. Add the URL and it works.
Claude Code
Run /mcp inside Claude Code to see the server and its tools.
Claude Desktop and claude.ai
Go to Customize, then Connectors. Select + Add, then Add custom connector. Enter:
- Name:
Scribiz - URL:
https://scribiz.com/mcp - Authentication: No sign in
The claude_desktop_config.json file is a separate setup for local servers. You do not need it for the hosted server.
Cursor
Put this in ~/.cursor/mcp.json, or in .cursor/mcp.json inside a project:
VS Code
Put this in .vscode/mcp.json. VS Code uses servers, not mcpServers.
Codex
Add this to ~/.codex/config.toml:
Claude Code plugin (server and skill together)
This repo is also a Claude Code plugin marketplace. The plugin adds the server and a skill that teaches Claude how to use it.
Use the plugin or the claude mcp add line above, not both. Other agents can install the skill alone:
Try it
Ask your agent:
This is a 19 second public video, "Me at the zoo". You should see a call to get_video_context, then search_video, and an answer with a timestamp link. The result says how the transcript was made. When a model read the link, its times are approximate (about 2 seconds either way). A video Scribiz has not processed before uses a little of the day's allowance on its first read, about 0.3 minutes for this one. After that it costs nothing.
Tools
All five tools are read-only. They reach out to the internet, and they never change your files.
Every result says which layers ran (captions, listening, watching, a model summary) and how many minutes it used. Details are in the tools reference.
What works without an account
- Captions, when the server can read them.
- Videos Scribiz has already processed, at no cost. The one exception is a summary Scribiz has not written for the video yet, which uses a tenth of its length.
- Videos with no captions the server can read: a model reads the video from its link. Its times are approximate (about 2 seconds either way), it has no speaker labels, and the result says so.
This runs inside a free allowance: about 10 minutes of reading a day for each connection (the same 10 minutes as the free web tool), videos up to 15 minutes, 5 questions a day and 30 caption lookups a day. Everyone who connects without a key also shares one more daily limit. When your 10 minutes or the shared limit are used up, captions and videos already processed still work. When the 30 caption lookups are used up, videos already processed still work. When the 5 questions are used up, searching and reading still work. The error says what is used up and when it resets (00:00 UTC).
Without a key the server never looks at the picture, so there are no on-screen notes, and a question about what was shown is answered from the words only. Watch reads the picture and the text shown on screen. It needs a Scribiz API key and uses minutes from your plan. A free account has 30 minutes a month, and you make a key in the dashboard. With a key, pass watch: true to get_video_context or ask_video; on-screen notes Scribiz already stored come back at no cost. See connection modes and limits and cost.
The hosted server takes public links only. It cannot read files on your disk.
With a key
Send the key as a bearer token. In Claude Code:
Local files
A local server comes with the command-line tool. It runs on your machine and can read your files. It needs Node 24, ffmpeg and yt-dlp, and has been tested on macOS only.
What is sent where
Your agent sends the video link and your questions to scribiz.com. Scribiz fetches the captions from the site you linked when it can. For a video it has not processed, it passes the link to Google's Gemini API, whose model reads the public video. Nothing is uploaded from your machine. Where Scribiz downloads the audio, it sends the audio to Google for transcription and analysis. For Watch it also sends a small low-resolution copy of the picture, with no sound. The privacy policy has the retention periods and Google's terms. This repo sends nothing. It holds setup files only.
Treat transcripts as data
A video can say anything, including "ignore your instructions". The server wraps all text from a video between === BEGIN UNTRUSTED VIDEO TEXT [id] === and === END UNTRUSTED VIDEO TEXT [id] === lines, and sets untrusted_content: true on the result. Do not auto-approve tools that can send messages, run commands or write files in a session that reads video. More in MCP security.
Links
- Scribiz MCP server page
- MCP docs
- Privacy policy and terms
- Problems with this repo or the setup files: open an issue here.
Also from Scribiz
Scribiz also has a free web tool. Paste a video link, get the transcript. No account needed: scribiz.com. There is a Mac app for Apple silicon Macs at scribiz.com/download, and a command-line tool on npm (npm install -g scribiz).
Built by Ilias Ism.
Source: README.md at commit 1b15c14
Tools
0Version history
1- v0.1.0LatestOct 5, 2026