
Gemini Media
io.github.mordor-forgev1.0.0Updated Oct 8, 2026
Generate images (Nano Banana), video (Veo, Gemini Omni), speech and music (Lyria) with Google models
Overview
Lets an assistant generate and edit images, video, speech and music with Google's generative media models, with cost estimates and budgets.
- What it does
- Wraps Google's generative media models for image generation and editing, video generation, extension and editing, text-to-speech and music creation. Tools include generate_image, edit_image, generate_video, get_video, extend_video, edit_video, generate_speech, generate_music, list_models, estimate_cost, get_usage and get_config. Results report the saved file path, a gemini-media://files/ URI that can be chained into other tools, and an estimated cost.
- When to use it
- Useful when an assistant needs to produce or modify media assets such as images, clips, voiceovers or music, including multi-asset production workflows. Also suited to cost-sensitive use, since it estimates spend before running a request and enforces session, daily and monthly budgets.
- Requirements
- A local binary on PATH, or the Docker image, Claude Desktop bundle or Gemini CLI extension. Credentials are either a Gemini API key (GEMINI_API_KEY) or a Google Cloud project with Vertex AI enabled (GOOGLE_CLOUD_PROJECT) plus Application Default Credentials. Network access to Google APIs is required; media is written to an output directory.
Installation
In SourceWeft
- Open Gemini Media in the dashboard and add it to a workspace.
- Enable the server for the chats that should use its tools.
Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.
Other MCP clients
Follow the launch instructions in the repository.
README
gemini-media-mcp
An MCP server for Google's generative media models: images (Nano Banana 2.1 / Pro), video (Veo 3.1 and Gemini Omni Flash), speech (Gemini 3.8 TTS) and music (Lyria 3.5). It ships as a single Go binary, speaks stdio and Streamable HTTP (MCP 2026-07-28), works with the Gemini API or Vertex AI, and comes with agent skills and plugin packaging for Claude Code, Codex, Gemini CLI, VS Code/Copilot, Cursor and more.
- Current models, updated without a release. A built-in catalog records IDs, aliases, lifecycle, parameters and prices. Retired models redirect to their replacement, and new model IDs work before the catalog knows them. You can override or extend the catalog with a hot-reloaded YAML file.
- Cost-aware. Every result reports its estimated cost, and
estimate_costcompares options before you spend. Spend is recorded in a ledger, capped by session, daily and monthly budgets, and calls above a threshold need explicit approval. - Agent-friendly. Each tool matches a workflow, and errors come back as
[kind] message + Hint. Image results include inline previews, and outputs can be chained by URI. Video runs as async jobs with long-polling and progress notifications. - Robust. Backend and credentials are detected the way the Google SDKs do it. Vertex locations are chosen per model, retries and timeouts are built in, files are written atomically with provenance, and HTTP mode ships with security defaults.
Upgrading from v0? The tools changed. See the migration table. The review also covers what was broken, why, and the design of v1.
Quick start
-
Get credentials. Either:
- an API key from Google AI Studio (
GEMINI_API_KEY), or - a Google Cloud project with Vertex AI enabled (
GOOGLE_CLOUD_PROJECTplusgcloud auth application-default login).
- an API key from Google AI Studio (
-
Install the binary. On macOS or Linux:
On Windows, download the
windowszip from the latest release and putgemini-media-mcp.exeon your PATH. With Go 1.26+:go install github.com/mordor-forge/gemini-media-mcp/cmd/gemini-media-mcp@latest.The Gemini CLI extension, the Claude Desktop bundle and the Docker image include the binary, so they skip this step.
-
Add the server to your agent. Every config below runs
gemini-media-mcpfrom your PATH.Claude Code (plugin: server plus skills):
or just the server:
claude mcp add gemini-media -e GEMINI_API_KEY=... -- gemini-media-mcpCodex (
~/.codex/config.toml):Gemini CLI (paid API keys, Vertex AI and Gemini Code Assist only, since June 2026):
gemini extensions install https://github.com/mordor-forge/gemini-media-mcpAntigravity CLI: add the server to
~/.gemini/config/mcp_config.json, as shown in the install snippets.Claude Desktop: download
gemini-media-mcp-<version>.mcpbfrom the latest release and open it.VS Code / Copilot:
code --add-mcp '{"name":"gemini-media","command":"gemini-media-mcp"}'Cursor, Windsurf and other
mcpServers-style clients (Zed, OpenCode and Goose use their own formats, see below):If a desktop app reports that
gemini-media-mcpwas not found, it doesn't see your shell's PATH: use the full path fromcommand -v gemini-media-mcp. Every client, plus Docker and the skills installer, is covered in packaging/INSTALL-SNIPPETS.md. -
If your agent doesn't forward environment variables (plugins often don't), store the key once:
To run it in Docker instead: docker run -i --rm --user "$(id -u):$(id -g)" -e GEMINI_API_KEY -v "$PWD/media:/output" -v gemini-media-state:/state ghcr.io/mordor-forge/gemini-media-mcp
(--user makes the generated files yours on Linux; the named /state volume keeps spend accounting and video jobs between runs).
Tools
Every result includes the saved file's path, a gemini-media://files/<name> URI and its cost. You can pass the URI as an input to another tool. Clients that can't read the server's disk (for example over HTTP) can fetch the file with resources/read.
Models
Use an alias or a full model ID. Run gemini-media-mcp models for the built-in catalog with prices, or call list_models, which also applies your override file and, with live: true, checks what your key can use.
Updating or adding a model without waiting for a release. Create a YAML file and point GEMINI_MEDIA_CATALOG at it. The server merges it with the built-in catalog by id and reloads it automatically:
See internal/catalog/models.yaml for the full schema.
Configuration
Settings are layered: built-in defaults < config file < environment < flags. The config file lives at ~/.config/gemini-media-mcp/config.yaml on Linux, ~/Library/Application Support/gemini-media-mcp/config.yaml on macOS, %AppData%\gemini-media-mcp\config.yaml on Windows, or wherever GEMINI_MEDIA_CONFIG points. Unknown keys produce warnings instead of errors, and get_config shows where each setting came from.
How the backend is chosen:
- An explicit
backendsetting wins. - Otherwise the SDK switches
GOOGLE_GENAI_USE_ENTERPRISEandGOOGLE_GENAI_USE_VERTEXAIdecide. - Otherwise an API key selects the Gemini API, even if
GOOGLE_CLOUD_PROJECTis set in your shell for other tools. - Otherwise a project selects Vertex AI.
On Vertex, an API key without a project uses express mode, which does not support video.
Spend and budgets
Google doesn't return costs, so the server computes them from its price table and the token usage the API reports. Every call records an estimate before it runs and the reconciled cost afterwards. Failed and safety-blocked generations count as $0. A response that used tokens but returned no media (for example MAX_TOKENS) is charged for those tokens.
- Results include
cost {estimatedUsd, usd, basis}.get_usageandgemini-media-mcp usagesummarize spend. The ledger is a plain JSONL file in the state directory. - Budgets are enforced before each call, across concurrent calls and across several server processes sharing a state directory (in-flight calls hold their estimate under a file lock).
- With
GEMINI_MEDIA_CONFIRM_ABOVE_USD=1, an expensive call (for example a $3.20 Veo clip) comes back as a[confirmation]error. The agent asks you, then retries withapprovedCostUsd. - These figures are estimates. For a hard stop, also set a spend cap in AI Studio or a budget in Cloud Billing.
HTTP mode
- HTTP mode is stateless Streamable HTTP, which MCP 2026-07-28 requires; older clients still work.
- The server protects against DNS rebinding and cross-origin requests.
- It refuses to listen on a non-loopback address without a token.
- It only reads input files from the output directory or
GEMINI_MEDIA_INPUT_DIRS(unlessGEMINI_MEDIA_ALLOW_ANY_INPUT_PATH=true). GET /healthzreports status.
Skills
The skills/ directory contains Agent Skills that teach agents the full workflow for each media type: intent, prompt craft, model choice, cost checks, review and iteration. They cover interactive use as well as unattended runs.
Plugin installs (Claude Code, Codex, Gemini CLI, VS Code) include the skills. To install them in any Agent Skills–compatible agent, run npx skills add mordor-forge/gemini-media-mcp, or copy the folders into .agents/skills/ or ~/.claude/skills/.
Development
Paid live E2E tests are run locally with your own API key. CI runs the free
checks and compiles/vets the E2E tests, but never executes live generation,
including after merges to main or on manual workflow runs.
- AGENTS.md: guide for coding agents and contributors.
- docs/architecture-review.md: architecture and design decisions.
- internal/catalog/models.yaml: model catalog. When Google changes models, edit this file, not the code.
License
Source: README.md at commit 747f3b8
Tools
0Version history
1- v1.0.0LatestOct 8, 2026


