Audio
io.github.audiojsv2.10.0Updated Oct 3, 2026
Edit, analyze and convert audio: loudness, spec checks, denoise, EQ, cuts, BPM, key. No ffmpeg.
Overview
Lets an assistant edit, analyze, and convert audio files locally — loudness, EQ, denoise, cuts, BPM and key — without ffmpeg.
- What it does
- Exposes a JavaScript audio library as MCP tools for decoding audio from files, URLs, bytes, or streams, then editing it: trim, shrink, crop, insert, cut, paste, move, split, reverse, speed, stretch, pitch, formant, and remix. It also processes audio with gain, fades, normalization to loudness targets, mixing, crossfades, panning, filters, EQ, denoise, vocal separation, and resampling. Analysis covers loudness, RMS, peak, noise floor, BPM and beat detection, and pass/fail checks against delivery specs such as ACX, podcast, streaming, broadcast, and Netflix. Results can be saved or encoded to formats like mp3, wav, flac, and m4a, or exported as cut lists for video editors.
- When to use it
- Useful when an assistant needs to prepare or inspect audio locally: cleaning up narration or podcast recordings, meeting loudness or delivery specs, trimming silence, converting formats, or extracting tempo and key. It suits workflows that would otherwise require ffmpeg or a DAW.
- Requirements
- Runs as a local process over stdio, installed from the npm package audio. It is desktop only and not available as a web executable. No authentication, environment variables, or headers are declared. Some optional features download model weights or need extra packages.
Installation
In SourceWeft
- Open Audio in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
audio [test] [npm]
Audio playback, editing and analysis
- Any Format — fast wasm codecs, no ffmpeg.
- Non-destructive — virtual edits, infinite undo, instant clone.
- Stream-first — playback/encode during decode, realtime editing.
- Paged — no 2Gb memory limit, open 10Gb+ files.
- Analysis — loudness, spectrum, beats, pitch, chords, key.
- Modular – pluggable ops, tree-shakable.
- CLI — playback, batch processing, scripting, unix pipes, tab completion.
- Cross-platform — browsers, node, deno, bun.
Start
Node
npm i audio
Browser
CLI
Skill
MCP
Prompt: make ~/Desktop/interview.m4a podcast-ready and tell me the loudness before and after
One agent at a time: claude mcp add audio -- npx -y audio --mcp, qwen mcp add audio npx -y audio --mcp, droid mcp add audio "npx -y audio --mcp".
Editor and agents
Connect the editor to it with the key (once: the bridge keeps it), and its chat runs an agent of yours, which measures, looks at, edits and plays the sound open there; each tab keeps its conversations, each with the agent picked under the message. The bridge finds Claude Code, Codex, Pi, Gemini CLI, Qwen Code, Kimi Code, OpenCode, Kilo Code, Cline, Goose, Factory Droid, Cursor, Augment, Kiro and Mistral Vibe on PATH; any other that speaks ACP runs by its command line, --agent "my-agent --acp".
Any MCP agent gets the editor's tools (state, measure, look, edit, select, play, check, …), the running bridge found by itself: npx add-mcp "npx -y audio --mcp --editor".
An agent thinks with the model its own settings name: Pi takes local ones for good from ollama launch pi --config; Claude Code takes any Anthropic-compatible endpoint from the bridge's environment:
Recipes
Clean up
Master & deliver
Compose
Analyze
Record & generate
Automate
Stream & persist
API
Create
Properties
Structure
Every op takes a trailing {at, duration, channel} range, except channel-changing remix and crossover. Times are seconds or strings ('1:30', '2m'); negative counts from the end. FFmpeg's short names work wherever the long ones do: d for duration, xfade for crossfade (a.remove({ at: 1, d: 0.5, xfade: 0.01 })).
Process
Filter
All biquads.
Effect
I/O
Playback / Recording
Playback renders up to 2 s ahead into an AudioWorklet on audio.context (Node: @audio/speaker), so a busy main thread doesn't stop it, and sounds within milliseconds of play() (the device's own latency aside). An edit to the playing instance is heard ~50 ms later where it happens: the audio rendered ahead gives way, crossfaded. A source still arriving (decoding, pushed) plays what has come and goes on as more comes. Any channel count plays as it is; the device downmixes.
Metering
Analysis
Opts: bpm, beats, onsets take { minBpm, maxBpm, delta, frameSize, hopSize }; notes takes { minFreq, maxFreq, frameSize, hopSize, minDuration }; chords, key take { frameSize, hopSize, tuning } (frames of 16384 samples at 44.1 kHz, 0.34 to 0.51 s at other rates, every eighth of a frame; concert A read from the audio unless tuning in Hz is given); chords also boostN (no-chord bias, 0.1); key also method: 'nnls' | 'pcp'. chords needs @audio/mir-nnls-chroma and @audio/mir-chordino, key needs @audio/mir-nnls-chroma: GPL-2.0-or-later translations of the reference plugins, installed by choice (npm i @audio/mir-nnls-chroma @audio/mir-chordino); key with method: 'pcp' needs only the MIT @audio/mir-chroma and @audio/mir-key, installed with audio unless optional dependencies are skipped. notes with robust: true needs @audio/neural-pitch (weights inside): a network's pitch candidates in place of YIN's keep the notes where YIN loses them (Vocadito onsets F 0.76 against 0.53 at 0 dB SNR) and trail it slightly on clean audio, so YIN stays the default. notes with poly: true takes { minFreq, maxFreq, minDuration, onsetThreshold, frameThreshold } and needs @audio/neural-transcribe, whose model downloads on first use; bends are cents from the note's pitch per 11.6 ms frame, in 33.3-cent steps (in-tune notes read 0).
Meta
Parsed on decode, written on save; round-trips WAV, MP3, FLAC.
Utility
Plugins
Plugins also run without the engine: audio/batch over a whole signal, audio/stream over live chunks. Plugin tutorial.
Worker
Across the boundary clip(), split(), clone() return promises; op errors emit 'error'; functions don't cross, use {t, v} curves. play() renders in the worker straight into the page's AudioWorklet: the main thread can stall for seconds without a dropout. Architecture.
CLI
npm i -g audio, or without installing: npx audio …
A pipeline: a source produces audio, transforms reshape it, a sink consumes it. The default sink is stat — printing an overview.
Playback
[Audiojs demo]␣ pause · ←/→ seek ±10s · ⇧←/⇧→ seek ±60s · ↑/↓ volume · l loop · s save as · q quit
Edit
Analysis
Batch
Stdin/stdout
Tab completion
FAQ
- What formats are supported?
- Decode: WAV, MP3, FLAC, OGG Vorbis, Opus, AAC, AIFF, CAF, WebM, AMR, WMA, QOA via decode. Encode: WAV, MP3, FLAC, Opus, OGG, AIFF via encode. Codecs are WASM-based, lazy-loaded on first use.
- Does it need ffmpeg or native addons?
- No, pure JS + WASM. For CLI, you can install globally:
npm i -g audio. - How big is the bundle?
- ~20K gzipped core. Codecs load on demand via
import(), so unused formats aren't fetched. - How does it handle large files?
- Audio is stored in fixed-size pages. In the browser, cold pages can evict to OPFS when memory exceeds budget — auto-sized from
navigator.storage.estimate()(quota/4, 64MB..2GB), overridable via{budget}. Stats stay resident (~7 MB for 2h stereo). - Are edits destructive?
- No.
a.gain(-3).trim()pushes entries to an edit list — source pages aren't touched. Edits replay onread()/save()/for await. - Can I use it in the browser?
- Yes, same API. See Browser for bundle options and import maps.
- Does it need the full file before I can work with it?
- No. Playback, edits, and structural ops (crop, repeat, pad, insert, etc.) all stream incrementally during decode — output begins before the file finishes loading. The edit plan recompiles as data arrives, tracking a safe output boundary per op. Only ops that depend on total length (open-end reverse, negative
at) wait for full decode. - TypeScript?
- Yes, ships with
audio.d.ts. - Does it have feature parity with FFmpeg / SoX / librosa?
- Yes — the audiojs ecosystem covers the practical baseline of FFmpeg filters, SoX effects, librosa analysis, Pedalboard and MIREX, all as
@audio/*pluginsaudiowires through one API (the few uncovered items are esoteric or deliberately skipped). Every effect, filter, generator and analyzer lives in the registry — call one by name and it loads on first use. Coverage matrix: docs/comparison.md. - How is this different from SoX / FFmpeg / Audacity / librosa / Web Audio / Tone.js?
- In one line:
audiois the only one that runs the same API in Node and the browser, with non-destructive lazy edits that stream during decode. The native tools (SoX, FFmpeg) are faster on raw throughput but have no JS API, browser, or undo; the browser libs (Web Audio, Tone.js, Howler) are real-time graphs, not file editors. Full feature and performance matrices vs pydub, librosa, aubio, essentia, Pedalboard, SoX, FFmpeg, Audacity and MATLAB are in docs/comparison.md.
Built with
- decode – codec decoding (13+ formats)
- encode – codec encoding
- filter – filters (weighting, EQ, auditory)
- speaker – audio output
- mic – audio input
- pitch – pitch, chord, key analysis
- audio-type – format detection
- pcm-convert – PCM format conversion
Source: README.md at commit c81e504
Tools
0Version history
1- v2.10.0LatestOct 3, 2026


