Audio
io.github.audiojsv2.10.0更新於 Oct 3, 2026
Edit, analyze and convert audio: loudness, spec checks, denoise, EQ, cuts, BPM, key. No ffmpeg.
概覽
讓助手在本機編輯、分析與轉換音訊——響度、EQ、降噪、剪輯、BPM 與調性——不需 ffmpeg。
- 功能
- 把一套 JavaScript 音訊函式庫包成 MCP 工具,可從檔案、URL、位元組或串流解碼音訊,再進行編輯:trim、shrink、crop、insert、cut、paste、move、split、reverse、speed、stretch、pitch、formant 與 remix。處理功能包括增益、淡入淡出、依響度目標正規化、混音、交叉淡化、聲像、濾波、EQ、降噪、人聲分離與重新取樣。分析涵蓋響度、RMS、峰值、噪聲底、BPM 與節拍偵測,以及依 ACX、播客、串流、廣播與 Netflix 等交付規範做通過/不通過檢查。結果可儲存或編碼為 mp3、wav、flac、m4a 等格式,也可匯出成剪輯軟體用的剪裁表。
- 適用情境
- 適合助手需要在本機準備或檢查音訊的情境:清理旁白或播客錄音、符合響度或交付規範、去除靜音、轉換格式,或擷取速度與調性。適用於原本需要 ffmpeg 或數位音訊工作站的工作流程。
- 執行需求
- 以本機程序透過 stdio 執行,從 npm 套件 audio 安裝。僅支援桌面端,不提供網頁可執行版本。未宣告驗證、環境變數或標頭。部分選用功能會下載模型權重或需要額外套件。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Audio,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
audio [test] [npm]
Audio playback, editing and analysis
- Any Format — fast wasm codecs, no ffmpeg.
- Non-destructive — virtual edits, infinite undo, instant clone.
- Stream-first — playback/encode during decode, realtime editing.
- Paged — no 2Gb memory limit, open 10Gb+ files.
- Analysis — loudness, spectrum, beats, pitch, chords, key.
- Modular – pluggable ops, tree-shakable.
- CLI — playback, batch processing, scripting, unix pipes, tab completion.
- Cross-platform — browsers, node, deno, bun.
Start
Node
npm i audio
Browser
CLI
Skill
MCP
Prompt: make ~/Desktop/interview.m4a podcast-ready and tell me the loudness before and after
One agent at a time: claude mcp add audio -- npx -y audio --mcp, qwen mcp add audio npx -y audio --mcp, droid mcp add audio "npx -y audio --mcp".
Editor and agents
Connect the editor to it with the key (once: the bridge keeps it), and its chat runs an agent of yours, which measures, looks at, edits and plays the sound open there; each tab keeps its conversations, each with the agent picked under the message. The bridge finds Claude Code, Codex, Pi, Gemini CLI, Qwen Code, Kimi Code, OpenCode, Kilo Code, Cline, Goose, Factory Droid, Cursor, Augment, Kiro and Mistral Vibe on PATH; any other that speaks ACP runs by its command line, --agent "my-agent --acp".
Any MCP agent gets the editor's tools (state, measure, look, edit, select, play, check, …), the running bridge found by itself: npx add-mcp "npx -y audio --mcp --editor".
An agent thinks with the model its own settings name: Pi takes local ones for good from ollama launch pi --config; Claude Code takes any Anthropic-compatible endpoint from the bridge's environment:
Recipes
Clean up
Master & deliver
Compose
Analyze
Record & generate
Automate
Stream & persist
API
Create
Properties
Structure
Every op takes a trailing {at, duration, channel} range, except channel-changing remix and crossover. Times are seconds or strings ('1:30', '2m'); negative counts from the end. FFmpeg's short names work wherever the long ones do: d for duration, xfade for crossfade (a.remove({ at: 1, d: 0.5, xfade: 0.01 })).
Process
Filter
All biquads.
Effect
I/O
Playback / Recording
Playback renders up to 2 s ahead into an AudioWorklet on audio.context (Node: @audio/speaker), so a busy main thread doesn't stop it, and sounds within milliseconds of play() (the device's own latency aside). An edit to the playing instance is heard ~50 ms later where it happens: the audio rendered ahead gives way, crossfaded. A source still arriving (decoding, pushed) plays what has come and goes on as more comes. Any channel count plays as it is; the device downmixes.
Metering
Analysis
Opts: bpm, beats, onsets take { minBpm, maxBpm, delta, frameSize, hopSize }; notes takes { minFreq, maxFreq, frameSize, hopSize, minDuration }; chords, key take { frameSize, hopSize, tuning } (frames of 16384 samples at 44.1 kHz, 0.34 to 0.51 s at other rates, every eighth of a frame; concert A read from the audio unless tuning in Hz is given); chords also boostN (no-chord bias, 0.1); key also method: 'nnls' | 'pcp'. chords needs @audio/mir-nnls-chroma and @audio/mir-chordino, key needs @audio/mir-nnls-chroma: GPL-2.0-or-later translations of the reference plugins, installed by choice (npm i @audio/mir-nnls-chroma @audio/mir-chordino); key with method: 'pcp' needs only the MIT @audio/mir-chroma and @audio/mir-key, installed with audio unless optional dependencies are skipped. notes with robust: true needs @audio/neural-pitch (weights inside): a network's pitch candidates in place of YIN's keep the notes where YIN loses them (Vocadito onsets F 0.76 against 0.53 at 0 dB SNR) and trail it slightly on clean audio, so YIN stays the default. notes with poly: true takes { minFreq, maxFreq, minDuration, onsetThreshold, frameThreshold } and needs @audio/neural-transcribe, whose model downloads on first use; bends are cents from the note's pitch per 11.6 ms frame, in 33.3-cent steps (in-tune notes read 0).
Meta
Parsed on decode, written on save; round-trips WAV, MP3, FLAC.
Utility
Plugins
Plugins also run without the engine: audio/batch over a whole signal, audio/stream over live chunks. Plugin tutorial.
Worker
Across the boundary clip(), split(), clone() return promises; op errors emit 'error'; functions don't cross, use {t, v} curves. play() renders in the worker straight into the page's AudioWorklet: the main thread can stall for seconds without a dropout. Architecture.
CLI
npm i -g audio, or without installing: npx audio …
A pipeline: a source produces audio, transforms reshape it, a sink consumes it. The default sink is stat — printing an overview.
Playback
[Audiojs demo]␣ pause · ←/→ seek ±10s · ⇧←/⇧→ seek ±60s · ↑/↓ volume · l loop · s save as · q quit
Edit
Analysis
Batch
Stdin/stdout
Tab completion
FAQ
- What formats are supported?
- Decode: WAV, MP3, FLAC, OGG Vorbis, Opus, AAC, AIFF, CAF, WebM, AMR, WMA, QOA via decode. Encode: WAV, MP3, FLAC, Opus, OGG, AIFF via encode. Codecs are WASM-based, lazy-loaded on first use.
- Does it need ffmpeg or native addons?
- No, pure JS + WASM. For CLI, you can install globally:
npm i -g audio. - How big is the bundle?
- ~20K gzipped core. Codecs load on demand via
import(), so unused formats aren't fetched. - How does it handle large files?
- Audio is stored in fixed-size pages. In the browser, cold pages can evict to OPFS when memory exceeds budget — auto-sized from
navigator.storage.estimate()(quota/4, 64MB..2GB), overridable via{budget}. Stats stay resident (~7 MB for 2h stereo). - Are edits destructive?
- No.
a.gain(-3).trim()pushes entries to an edit list — source pages aren't touched. Edits replay onread()/save()/for await. - Can I use it in the browser?
- Yes, same API. See Browser for bundle options and import maps.
- Does it need the full file before I can work with it?
- No. Playback, edits, and structural ops (crop, repeat, pad, insert, etc.) all stream incrementally during decode — output begins before the file finishes loading. The edit plan recompiles as data arrives, tracking a safe output boundary per op. Only ops that depend on total length (open-end reverse, negative
at) wait for full decode. - TypeScript?
- Yes, ships with
audio.d.ts. - Does it have feature parity with FFmpeg / SoX / librosa?
- Yes — the audiojs ecosystem covers the practical baseline of FFmpeg filters, SoX effects, librosa analysis, Pedalboard and MIREX, all as
@audio/*pluginsaudiowires through one API (the few uncovered items are esoteric or deliberately skipped). Every effect, filter, generator and analyzer lives in the registry — call one by name and it loads on first use. Coverage matrix: docs/comparison.md. - How is this different from SoX / FFmpeg / Audacity / librosa / Web Audio / Tone.js?
- In one line:
audiois the only one that runs the same API in Node and the browser, with non-destructive lazy edits that stream during decode. The native tools (SoX, FFmpeg) are faster on raw throughput but have no JS API, browser, or undo; the browser libs (Web Audio, Tone.js, Howler) are real-time graphs, not file editors. Full feature and performance matrices vs pydub, librosa, aubio, essentia, Pedalboard, SoX, FFmpeg, Audacity and MATLAB are in docs/comparison.md.
Built with
- decode – codec decoding (13+ formats)
- encode – codec encoding
- filter – filters (weighting, EQ, auditory)
- speaker – audio output
- mic – audio input
- pitch – pitch, chord, key analysis
- audio-type – format detection
- pcm-convert – PCM format conversion
來源:README.md,提交 c81e504
工具
0版本歷史
1- v2.10.0最新Oct 3, 2026


