babelscribe

io.github.phonology024v0.3.2更新於 Oct 6, 2026

Offline speech-to-text and SRT subtitles in 99 languages on your own AMD/NVIDIA/Intel/Apple GPU.

已驗證STDIO僅桌面Files & StorageMedia & Design

概覽

AI 產生的概覽

在自己的 GPU 或 CPU 上,離線把音訊與視訊檔轉成 99 種語言的字幕與文字。

功能
babelscribe 包裝 whisper.cpp,把音訊與視訊檔轉成 SRT、VTT、TXT 與 JSON 輸出,可自動偵測語言或指定語言。它的 MCP 工具包括 transcribe(檔案轉字幕加文字)、find_media(在常用資料夾中尋找最新的音訊或視訊)、list_devices 與 list_languages_and_models。它會自動挑選合適的 GPU,內建 ffmpeg 處理媒體轉換,並提供精確模式與部分語言的混合微調模型。
適用情境
適合希望助理從本機媒體產生字幕或轉錄文字、又不想把音訊上傳到雲端的情況,包括只有 AMD、Intel 或 Apple GPU,而 CUDA 專用工具只能退回 CPU 的環境。也適合一步找到最近的錄音並完成轉錄。
執行需求
以本機 stdio 程序執行,通常使用 uvx babelscribe mcp 或已安裝的 babelscribe 指令;pip 安裝需要 Python 3.9+。未宣告帳號、API 金鑰或環境變數。首次使用需要網路連線下載 whisper-cli 程式、語音模型與 Python 套件,並在 ~/.babelscribe 下佔用磁碟空間快取模型。
安裝前請注意
它會在媒體檔旁或指定資料夾寫入字幕與文字檔,並把轉錄文字回傳給呼叫它的 AI 應用,之後適用該應用自身的隱私政策。首次使用會從 GitHub、Hugging Face 與 PyPI 下載程式與模型,這些站點會看到一般下載請求。長影片可能超過 MCP 工具預設逾時,專有名詞也可能拼錯,發布字幕前請核對名稱。

安裝

在 SourceWeft 中

  1. 開啟 儀表板中的 babelscribe,將其新增到工作區。
  2. 為需要使用其工具的對話啟用該服務。

Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。

其他 MCP 客戶端

參照 儲存庫 中的啟動說明。

README

babelscribe — free, offline speech-to-text and subtitles on any GPU (AMD, NVIDIA, Intel, Apple)

[PyPI] [License: MIT] [MCP server]

babelscribe turns any audio or video file into subtitles (SRT, VTT) and text, in any of Whisper's 99 languages, on your own graphics card — AMD Radeon, NVIDIA, Intel Arc or Apple Silicon — or the CPU. No CUDA needed, no cloud, free (MIT). It runs OpenAI Whisper through whisper.cpp with Vulkan, CUDA or Metal, and works from the command line, from Python, or from AI apps (Claude, Codex, Antigravity, Gemini CLI, Cursor) as an MCP server.

bash
pip install babelscribe          # Python 3.9+babelscribe interview.mp4                 # auto-detects language, picks your best GPU -> interview.srt + interview.json

babelscribe is a thin, friendly layer on top of whisper.cpp. It exists because getting fast Whisper on a non-NVIDIA card (e.g. an AMD Radeon on Windows) still means building whisper.cpp with the Vulkan SDK yourself. babelscribe downloads a prebuilt whisper-cli for your system and handles the rest.

Your hardwareBackend usedPrebuilt asset
AMD / NVIDIA / Intel GPU on WindowsVulkanwindows-x64-vulkan
NVIDIA on WindowsVulkan (same build)windows-x64-vulkan
NVIDIA on Linux (alternative)CUDAlinux-x64-cuda (--flavor cuda)
AMD / NVIDIA / Intel GPU on LinuxVulkanlinux-x64-vulkan
Apple SiliconMetalmacos-arm64-metal
No usable GPUCPU*-cpu (--flavor cpu)

A 19-minute English TED talk transcribes in 48 seconds on an AMD Radeon RX 9070 XT with 1.5% word error — see the benchmark below.

Use it from an AI app (MCP) — no terminal needed

babelscribe is also an MCP server, so any AI app that supports local MCP servers can transcribe files on your computer with your GPU. Ask in plain words — "make Thai subtitles for the newest video in my Downloads" — and the app finds the file, transcribes it, writes .srt / .txt next to it, and can then proofread names, translate or summarise the text for you.

Tools: transcribe (file → subtitles + text), find_media (newest audio/video in Downloads, Videos, Desktop, …), list_devices, list_languages_and_models.

Claude Desktop (one click): download babelscribe.mcpb from the latest release and double-click it (or Settings → Extensions → Install extension).

Everything else runs the same command — uv fetches babelscribe for you: uvx babelscribe mcp

AppHow to add it
Claude Codeclaude mcp add babelscribe -- uvx babelscribe mcp
OpenAI Codex CLIcodex mcp add babelscribe -- uvx babelscribe mcp, then in ~/.codex/config.toml under [mcp_servers.babelscribe] set tool_timeout_sec = 3600 and startup_timeout_sec = 120 (defaults are 60 s / 10 s — too short for a long video or the first model download)
Google Antigravityagent panel → ⋯ → MCP Servers → Manage MCP Servers → View raw config, add the JSON below
Gemini CLIadd the JSON below to ~/.gemini/settings.json
Cursoradd the JSON below to ~/.cursor/mcp.json
VS Code (Copilot).vscode/mcp.json: same entry under "servers" with "type": "stdio"
json
{  "mcpServers": {    "babelscribe": { "command": "uvx", "args": ["babelscribe", "mcp"], "timeout": 3600000 }  }}

(timeout is in milliseconds and only some apps read it.) Already installed with pip? Use "command": "babelscribe", "args": ["mcp"]. Web-only chat apps (e.g. grok.com, chatgpt.com) can only reach servers on the internet, not your computer, so they can't use your GPU or files.

Benchmark: real talks, human captions as the answer key

Four TED / TEDx talks, scored against the human-made captions in the spoken language (bench/bench.py, reproducible). Error = word error rate (WER) for space-separated languages, character error rate (CER) for Japanese and Thai. GPU: AMD Radeon RX 9070 XT via Vulkan. default = large-v3-turbo; accurate = --accurate (see the FLEURS section).

LanguageTalkLengthdefault: time / error--accurate: time / error
EnglishMatt Walker — Sleep Is Your Superpower (TED)19.3 min48 s (24x) · WER 1.5%107 s (11x) · WER 1.5%
JapaneseKazunari Taguchi (TEDxHimi)16.2 min47 s (20x) · CER 4.7%115 s (8x) · CER 4.2%
SpanishAdrià Solà Pastor — Cómo hablar (TEDxESIC University)19.9 min56 s (21x) · WER 13.8%134 s (9x) · WER 13.5%
Thaiนิติ ชัยชิตาทร — โปรดเรียกฉันด้วยนามอันแท้จริง (TEDxBangkok)14.1 min74 s (12x) · CER 22.7%565 s (2x) · CER 16.4%

How to read it: TED captions are edited for reading (fillers dropped, light rewording), so these numbers are an upper bound — most of the Spanish "errors" are the speaker's actual words versus the tidied caption. Thai is genuinely harder: fast, casual speech with slang; --accurate (Pathumma Whisper text + turbo timing) cuts its error by more than a quarter. Talks are used only to measure accuracy; their transcripts are not redistributed (TED content is CC BY-NC-ND).

Benchmark: 16 languages on FLEURS

Google FLEURS test set, 50 utterances per language, human-verified verbatim transcripts (bench/fleurs.py, reproducible). Same normaliser for every language (Whisper's rule: lower-case, drop punctuation and non-spacing marks, numbers spelled out). WER for space-separated languages, CER for ja / zh / ko / th.

Languagedefault (turbo)--accuratemodel --accurate uses--accurate, edge-trimmed*
English6.1%5.7%large-v3, beam 55.6%
Spanish4.1%4.2%large-v3, beam 52.6%
French7.4%7.3%large-v3, beam 57.2%
German4.3%4.0%large-v3, beam 53.3%
Portuguese8.5%7.7%large-v3, beam 55.3%
Italian6.7%6.5%large-v3, beam 53.8%
Russian6.8%5.8%large-v3, beam 55.6%
Arabic11.0%10.6%large-v3, beam 59.7%
Hindi28.4%12.6%vasista22/whisper-hindi-large-v2 + turbo timing11.1%
Indonesian9.8%7.8%large-v3, beam 55.6%
Vietnamese10.8%9.1%large-v3, beam 59.1%
Turkish6.2%6.7%large-v3, beam 55.8%
Japanese (CER)6.4%5.7%large-v3, beam 54.6%
Chinese (CER)6.3%5.3%large-v3, beam 54.8%
Korean (CER)4.1%3.9%large-v3, beam 53.1%
Thai (CER)15.9%8.9%Pathumma Whisper (NECTEC) + turbo timing8.8%

* Some FLEURS clips contain more speech than their reference transcript, so a correct model is charged for words the reference leaves out. Edge-trimmed ignores extra words before the first / after the last reference word; both numbers are stored by bench/fleurs.py. 50 utterances per language means differences under ~0.5 points are noise.

Also measured and not used, because large-v3 was as good or better: large-v2 (all languages), PhoWhisper-large (vi 19.0%), whisper-large-v3 dialectal / code-switching Arabic fine-tunes (12.2% / 15.2%), Typhoon Whisper (th 11.9%), Thonburian Whisper (th 9.1% — kept as an option), Vaani Hindi (16.5%).

Why babelscribe (vs. what already exists)

GPU on AMD / IntelWindows, no build stepVideo in, subtitles outLong files don't loopBetter text for your language
babelscribe✅ Vulkan✅ prebuilt whisper-cli downloaded for you✅ ffmpeg bundled✅ --max-context 0 by default✅ hybrid fine-tune text + turbo timing
whisper.cpp (raw)✅ Vulkan — if you compile it❌ official releases ship no Windows Vulkan build❌ WAV 16 kHz only⚠️ you must know the flag❌
faster-whisper / WhisperX❌ GPU = NVIDIA CUDA only✅ pip✅⚠️⚠️ manual
Cloud APIsn/a (cloud)✅✅✅❌ — and your audio leaves your machine, paid per minute

babelscribe does not replace those projects — it stands on whisper.cpp and simply removes the hard parts: compiling for your GPU, converting media, picking the right device, avoiding the long-file repeat bug, and combining a language-specific fine-tune with accurate timestamps.

FAQ

How do I run Whisper on an AMD GPU on Windows? pip install babelscribe, then babelscribe video.mp4. It downloads a Vulkan build of whisper.cpp that runs on AMD Radeon (and Intel Arc / NVIDIA) cards on Windows and Linux — no ROCm, no CUDA, no compiling.

How do I make subtitles (SRT) from a video for free, offline? babelscribe video.mp4 -f srt writes video.srt next to the video. Nothing is uploaded; it runs on your own computer.

What is the most accurate free transcription for Thai? babelscribe video.mp4 -l th --accurate — Pathumma Whisper (NECTEC) for the text plus Whisper turbo for timing: CER 8.9% on Google FLEURS vs 15.9% for plain Whisper turbo. Thonburian Whisper is available too (--text-model thai-thonburian).

Can Claude / ChatGPT Codex / Gemini transcribe a video on my computer? Yes — add babelscribe as an MCP server (see Use it from an AI app). Claude Desktop installs it with one click from babelscribe.mcpb. The AI app can then find a file, transcribe it on your GPU, and proofread, translate or summarise the text.

How fast is it? About 20x real time with Whisper large-v3-turbo on an AMD Radeon RX 9070 XT: a 19-minute talk in 48 seconds.

Which languages are supported? All 99 Whisper languages, with automatic language detection. FLEURS error rates for 16 of them are in the benchmark table.

Is it better than faster-whisper or WhisperX? Those are excellent on NVIDIA GPUs; on AMD / Intel GPUs they fall back to the CPU. babelscribe's niche is any GPU, zero setup, and better Thai / Hindi through community fine-tunes. If you have an NVIDIA card and like Python, faster-whisper is a fine choice.

Languages

All 99 languages Whisper was trained on, auto-detected or forced with -l: af am ar as az ba be bg bn bo br bs ca cs cy da de el en es et eu fa fi fo fr gl gu ha haw he hi hr ht hu hy id is it ja jw ka kk km kn ko la lb ln lo lt lv mg mi mk ml mn mr ms mt my ne nl nn no oc pa pl ps pt ro ru sa sd si sk sl sn so sq sr su sv sw ta te tg th tk tl tr tt uk ur uz vi yi yo yue zh

Accuracy follows Whisper's own training data: excellent for high-resource languages (English, Spanish, Japanese, …), weaker for low-resource ones. That is what hybrid mode is for — a community fine-tune for one language can be plugged in with one line in babelscribe/models.py (Thai and Hindi so far). PRs adding fine-tunes for other languages are the most valuable contribution.

Limitations (honest)

  • Tested end to end so far on an AMD Radeon RX 9070 XT (Windows, Vulkan). CUDA, Linux and macOS builds are produced by CI; reports from those machines are welcome.
  • Hybrid mode runs two models, so it is slower (≈1 min per minute of audio with a large fine-tune on that GPU).
  • Proper nouns can still be misspelled — check names before publishing subtitles.

Features

  • Any input — mp4, mkv, mov, mp3, wav, m4a… (ffmpeg is bundled through imageio-ffmpeg).
  • Any language — -l auto detects it; or pass -l th, -l ja, -l es…
  • Picks the right GPU — prefers a discrete card over an integrated one; babelscribe devices lists them, --device N overrides.
  • Long files that don't loop — runs whisper with --max-context 0, which stops the classic "same sentence repeated forever" hallucination on long recordings.
  • Hybrid mode for better spelling in your language — community fine-tunes (e.g. Thai Thonburian Whisper) spell far better but often lose timestamps. --text-model takes the text from the fine-tune and the timing from large-v3-turbo, aligned character by character (works for languages without spaces), cut only at word boundaries, with the timing model filling any words the fine-tune skipped.
  • --accurate — slower, fewest errors: large-v3 with beam search, or for Thai and Hindi the best community fine-tune (hybrid). Chosen per language from the FLEURS benchmark above.
  • Outputs — srt, vtt, txt, json (segments with token timings).

Usage

bash
babelscribe talk.mp4 -f srt,vtt,txt,json          # all formatsbabelscribe podcast.mp3 -l en -m large-v3          # pick language and modelbabelscribe talk.mp4 -l hi --accurate              # slower, fewest errors (best model per language)babelscribe vo.wav -l th --text-model thai-thonburian   # hybrid: pick a Thai fine-tune yourselfbabelscribe devices                                # GPUs whisper.cpp can seebabelscribe models                                 # models and fine-tunesbabelscribe talk.mp4 --bin /path/to/whisper-cli    # use your own whisper.cpp build

From Python:

python
from babelscribe.api import transcribe_filer = transcribe_file("talk.mp4", lang="auto", formats=["srt", "txt"])   # or accurate=Trueprint(r["lang"], r["device"], r["files"]); print(r["text"][:200])

Models download on first use to ~/.babelscribe/models (BABELSCRIBE_MODELS to change). Hybrid fine-tunes are converted on your machine from their original Hugging Face repo — install pip install "babelscribe[finetune]" once; converted weights are never redistributed, so each fine-tune keeps its own licence. Thai word boundaries: pip install "babelscribe[thai]".

Add a language fine-tune

Add one entry to FINETUNES in babelscribe/models.py (Hugging Face repo, language code, whether it needs beam search) and open a PR.

Prebuilt binaries

.github/workflows/build-binaries.yml builds whisper-cli + whisper-quantize for every row of the table above and attaches them to each v* release. Point BABELSCRIBE_RELEASES at another URL to self-host.

ภาษาไทย

ใช้ผ่าน Claude Desktop ได้โดยไม่ต้องพิมพ์คำสั่ง: โหลด babelscribe.mcpb จากหน้า Releases แล้วดับเบิลคลิก จากนั้นพิมพ์ในแชทว่า "ทำซับไทยให้คลิปล่าสุดในโฟลเดอร์ดาวน์โหลด" ถอดเสียงจากไฟล์เสียงหรือวิดีโอได้ทุกภาษา บนการ์ดจอทุกยี่ห้อ (AMD / NVIDIA / Intel ผ่าน Vulkan, Apple ผ่าน Metal) หรือ CPU ภาษาไทยแนะนำ babelscribe ไฟล์.mp4 -l th --accurate — ข้อความจาก Pathumma Whisper (NECTEC) ที่ผิดน้อยที่สุดใน FLEURS (CER 8.9% เทียบ turbo 15.9%) + เวลาจาก large-v3-turbo หรือเลือก Thonburian Whisper เอง: --text-model thai-thonburian

Privacy Policy

babelscribe runs entirely on your own computer. Last updated 2026-10-06.

  • Data collection: none. babelscribe has no telemetry, analytics, accounts or crash reporting, and never uploads your audio, video, transcripts or file names anywhere.
  • What it processes and where: the media file you choose is converted and transcribed locally; subtitles and text are written to your disk (next to the file, or the folder you choose). Through MCP, the transcript text is returned to the AI app that called the tool — what that app does with it is governed by that app's own privacy policy.
  • Network access (downloads only): on first use it downloads the whisper-cli program from this project's GitHub releases, speech models from Hugging Face (huggingface.co/ggerganov/whisper.cpp, and for --accurate Thai/Hindi the fine-tune's own repository), and Python packages from PyPI when installed with pip/uv. These are plain downloads; no personal data is sent. GitHub, Hugging Face and PyPI see a normal download request (IP address, user agent) under their own privacy policies.
  • Storage and retention: downloaded programs and models are cached in ~/.babelscribe (or BABELSCRIBE_HOME / BABELSCRIBE_MODELS) until you delete that folder. Outputs stay wherever they were written until you delete them. babelscribe keeps no other data.
  • Third-party sharing: none.
  • Contact: open an issue at https://github.com/phonology024/babelscribe/issues

Credits & licence

MIT. Built on whisper.cpp (MIT) and OpenAI Whisper models (MIT). Fine-tunes belong to their authors: Pathumma Whisper by NECTEC, Thonburian Whisper by biodatlab, whisper-hindi-large-v2 by vasista22 (Speech Lab, IIT Madras).

來源:README.md,提交 626383a

工具

0
工具後設資料尚未被收錄。

版本歷史

1
  1. v0.3.2最新Oct 6, 2026