
Audivo
io.github.AudivoDotDevv0.3.0更新於 Sep 30, 2026
Podcast transcripts: send an episode link and get the transcript back, a page at a time.
概覽
把播客單集轉成可分頁讀取的文字稿,並提供搜尋、報價以及本機音訊或 YouTube 上傳。
- 功能
- Audivo 為助理提供十個播客文字稿工具。transcribe 接受單集連結、帶 guid 的 feed_url、episode_id 或本機檔案,按頁回傳文字稿;read_transcript 繼續讀取後續頁面。search_shows、chart_shows、list_episodes、quote、confirm、group_status、list_groups 和 cancel_group 負責尋找節目與批次任務。本機伺服器另外提供 upload_audio 和 youtube_search。
- 適用情境
- 適合讓助理轉寫或摘要播客單集、搜尋節目與排行榜,或以先報價再確認的方式批次轉寫歷史節目。本機伺服器適合會啟動子程序的用戶端,並需要 YouTube 或本機檔案輸入;託管端點適合只接受 URL 的用戶端。
- 執行需求
- 託管方式:使用 streamable HTTP 端點,透過 OAuth 登入或 Authorization Bearer API 金鑰驗證。本機方式:需要 Node.js 與 npx,並設定 AUDIVO_API_KEY 環境變數。選用變數:AUDIVO_API_BASE_URL、AUDIVO_YTDLP_PATH、AUDIVO_CACHE_DIR。YouTube 支援會在首次使用時下載 yt-dlp。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Audivo,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Web executable,透過 Streamable HTTP。 遠端服務在工作區中設定後即可從網頁執行環境執行。
其他 MCP 客戶端
把它新增到你客戶端的 mcpServers 設定中。
{
"mcpServers": {
"mcp": {
"type": "http",
"url": "https://api.audivo.dev/mcp"
}
}
}README
Audivo MCP server
Podcast transcripts for Claude, ChatGPT, Codex, Cursor, and any other MCP client, backed by the Audivo API. Ask for an episode; get the transcript back.
Transcribe the latest episode of Acquired and summarise it.
Two ways to connect. The hosted endpoint serves ten tools. The local server serves the same ten, with
transcribe also taking a file on your machine or a YouTube link, plus upload_audio and
youtube_search.
Tools
transcribe
Pass one episode: an Apple Podcasts link, feed_url with guid, an episode_id from
list_episodes, or an upload_id — and on the local server, a YouTube link or an absolute path.
- An episode that is already transcribed comes back at once.
- A fresh one becomes a job.
transcribewaits for it inside the call: up to 20 seconds on the hosted server, whose gateway allows 29, and 50 by default on the local one, which reports progress to clients that ask for it. If it is still running, the answer is thejob_idandread_transcriptwaits the rest. - The transcript comes a page at a time, about 40,000 characters each (roughly 50 minutes of speech). Each page names the exact call for the next one.
max_creditsrefuses the call, before anything is spent, if it could cost more. A job holds its ceiling (the estimate plus 25%) and settles at the audio it measured, never above.- The same call twice returns the same job, never a second charge: the idempotency key is derived from the request.
For many episodes at once — a chart, a back catalogue — use quote and then confirm, which
refuses unless the model restates the quote's total.
Local: run it with npx
You need an API key from the Audivo dashboard. Keys are shown once. Set the key in the environment the client starts the server with, then add the server.
Claude Code (or install the Audivo plugin, which adds this server and the skill together)
Codex
Cursor, Claude Desktop, Windsurf, and other JSON-configured clients
VS Code (.vscode/mcp.json)
Keep files that contain a real key out of Git and shared chats.
Updating
@latest makes npx look the package up on npm each time the client starts the server, so a new
release arrives with your next session. To get it in the session you are in, restart the server: in
Claude Code, /mcp, choose audivo, then Reconnect. A spec without @latest works too, but
a bare @audivo/mcp resolves to a copy in the current project first, when there is one. The hosted
server at https://api.audivo.dev/mcp is kept current by Audivo.
Environment
The server writes nothing to stdout except protocol messages. Log lines go to stderr as JSON, and the key never appears in them.
YouTube (local server only)
When a show has no public RSS feed — many exist only on YouTube — search_shows finds nothing.
On the local server, youtube_search finds the episode and transcribe takes its link:
Find the Costco episode of Acquired on YouTube and transcribe it
transcribe downloads the video's audio on your machine with yt-dlp,
audio only and with no transcoding, uploads it to Audivo as your own upload, and transcribes it. The
transcript is private to your account.
yt-dlp is found in this order: AUDIVO_YTDLP_PATH; a yt-dlp on your PATH; otherwise the project's
official standalone binary for your platform, downloaded once from its GitHub releases into your user
cache directory and checked against that release's SHA2-256SUMS before it is made executable. The
package has no install script: nothing is downloaded until the first YouTube call. A managed copy
that stops working is updated to the latest release once and the call retried; a yt-dlp you
installed yourself is never touched. yt-dlp is given the Node running this server as its JavaScript
runtime.
You run the download, on your machine, under your own account; you are responsible for having the rights to transcribe what you download, as with any upload. Audivo's servers never fetch from YouTube, which is why the hosted server does not offer any of this.
Upload your own audio
A recording you made, an interview, or any file you already have on disk:
Transcribe /Users/alex/Downloads/interview.m4a
transcribe with path announces the file to Audivo (its hash, size, content type, and duration),
uploads the bytes straight to Audivo's storage with the signed URL the announcement returns, and
transcribes it. upload_audio does only the first two steps and returns an upload_id, for when you
want to quote several uploads together. The path must be absolute.
Limits: 1 byte to 5 GiB, up to 10 hours, and one of these content types: audio/mpeg, audio/mp3,
audio/mp4, audio/m4a, audio/x-m4a, audio/aac, audio/x-aac, audio/ogg, audio/opus,
audio/flac, audio/x-flac, audio/wav, audio/x-wav, audio/webm. An upload is kept for 7
days, and its transcript is private to your account; each account may hold up to 10 GiB across 100
unexpired uploads at a time. See the uploads guide.
Hosted: point a client at the URL
Clients that support MCP authorization — Claude, ChatGPT, Claude Code, VS Code — need only the URL: they open an Audivo sign-in page, you approve the connection, and it appears under Connected apps in the dashboard, where you can revoke it. For example, in Claude Code:
For clients that take a fixed header instead, or for automation, send an API key:
Per-client instructions are in the connection guide. The hosted
server is this package's lambda export, deployed by Audivo.
How it works
- Stateless. Every tool call is one or a few typed calls on the public API with the caller's own credential. The server keeps no table, no cache, and no copy of a credential beyond the call in flight.
- Fenced. Show names, episode and video titles, descriptions, and transcript text are publisher-authored. Each reaches the model inside a fence marked as untrusted content, with a per-response nonce, so a podcast cannot smuggle instructions into a session.
- Bounded.
transcribespends within the job's ceiling and the caller'smax_credits;confirmcompares the total the model states with the total the quote carried and refuses on a mismatch without sending anything. The API applies both checks on its side too. - Typed by the contract.
src/contract/types.tsis generated from the published OpenAPI spec incontract/openapi.yaml; a test fails the build when the two drift. - Local only.
upload_audio,youtube_search, and the file and YouTube inputs oftranscriberun on your machine; the hosted server never sees your files and never fetches from third-party sites.
Development
To pick up a spec change: npm run contract:sync fetches the published spec and regenerates the
types. Run it locally against a key with:
Releasing
Bump version in package.json, both version fields in server.json, and SERVER_INFO in
src/server.ts (a test fails until they agree), add a CHANGELOG.md entry, commit, then tag
v<version> and push the tag. The release workflow publishes to npm with provenance and lists the
release in the MCP Registry as
io.github.AudivoDotDev/mcp.
License
來源:README.md,提交 5940178
工具
0版本歷史
1- v0.3.0最新Sep 30, 2026
