
Gemini Media
io.github.mordor-forgev1.0.0更新於 Oct 8, 2026
Generate images (Nano Banana), video (Veo, Gemini Omni), speech and music (Lyria) with Google models
概覽
讓助理使用 Google 生成式媒體模型產生與編輯圖像、影片、語音和音樂,並提供費用估算與預算控制。
- 功能
- 封裝 Google 的生成式媒體模型,用於圖像生成與編輯、影片生成、延長與編輯、文字轉語音以及音樂創作。工具包括 generate_image、edit_image、generate_video、get_video、extend_video、edit_video、generate_speech、generate_music、list_models、estimate_cost、get_usage 和 get_config。結果會提供儲存的檔案路徑、可傳給其他工具的 gemini-media://files/ URI,以及估算費用。
- 適用情境
- 適合助理需要產出或修改媒體素材的情境,例如圖像、短片、配音或音樂,也支援多素材製作流程。由於會在執行前估算費用並執行工作階段、每日與每月預算,也適合對成本敏感的使用。
- 執行需求
- 需要 PATH 上的本機執行檔,或使用 Docker 映像、Claude Desktop 套件、Gemini CLI 擴充功能。憑證可以是 Gemini API 金鑰(GEMINI_API_KEY),或啟用 Vertex AI 的 Google Cloud 專案(GOOGLE_CLOUD_PROJECT)加應用程式預設憑證。需要連線至 Google API 的網路,媒體會寫入輸出目錄。
安裝
在 SourceWeft 中
- 開啟 儀表板中的 Gemini Media,將其新增到工作區。
- 為需要使用其工具的對話啟用該服務。
Desktop only,透過 STDIO。 STDIO 服務會啟動本機處理程序,因此需要 SourceWeft 桌面主機。
其他 MCP 客戶端
參照 儲存庫 中的啟動說明。
README
gemini-media-mcp
An MCP server for Google's generative media models: images (Nano Banana 2.1 / Pro), video (Veo 3.1 and Gemini Omni Flash), speech (Gemini 3.8 TTS) and music (Lyria 3.5). It ships as a single Go binary, speaks stdio and Streamable HTTP (MCP 2026-07-28), works with the Gemini API or Vertex AI, and comes with agent skills and plugin packaging for Claude Code, Codex, Gemini CLI, VS Code/Copilot, Cursor and more.
- Current models, updated without a release. A built-in catalog records IDs, aliases, lifecycle, parameters and prices. Retired models redirect to their replacement, and new model IDs work before the catalog knows them. You can override or extend the catalog with a hot-reloaded YAML file.
- Cost-aware. Every result reports its estimated cost, and
estimate_costcompares options before you spend. Spend is recorded in a ledger, capped by session, daily and monthly budgets, and calls above a threshold need explicit approval. - Agent-friendly. Each tool matches a workflow, and errors come back as
[kind] message + Hint. Image results include inline previews, and outputs can be chained by URI. Video runs as async jobs with long-polling and progress notifications. - Robust. Backend and credentials are detected the way the Google SDKs do it. Vertex locations are chosen per model, retries and timeouts are built in, files are written atomically with provenance, and HTTP mode ships with security defaults.
Upgrading from v0? The tools changed. See the migration table. The review also covers what was broken, why, and the design of v1.
Quick start
-
Get credentials. Either:
- an API key from Google AI Studio (
GEMINI_API_KEY), or - a Google Cloud project with Vertex AI enabled (
GOOGLE_CLOUD_PROJECTplusgcloud auth application-default login).
- an API key from Google AI Studio (
-
Install the binary. On macOS or Linux:
On Windows, download the
windowszip from the latest release and putgemini-media-mcp.exeon your PATH. With Go 1.26+:go install github.com/mordor-forge/gemini-media-mcp/cmd/gemini-media-mcp@latest.The Gemini CLI extension, the Claude Desktop bundle and the Docker image include the binary, so they skip this step.
-
Add the server to your agent. Every config below runs
gemini-media-mcpfrom your PATH.Claude Code (plugin: server plus skills):
or just the server:
claude mcp add gemini-media -e GEMINI_API_KEY=... -- gemini-media-mcpCodex (
~/.codex/config.toml):Gemini CLI (paid API keys, Vertex AI and Gemini Code Assist only, since June 2026):
gemini extensions install https://github.com/mordor-forge/gemini-media-mcpAntigravity CLI: add the server to
~/.gemini/config/mcp_config.json, as shown in the install snippets.Claude Desktop: download
gemini-media-mcp-<version>.mcpbfrom the latest release and open it.VS Code / Copilot:
code --add-mcp '{"name":"gemini-media","command":"gemini-media-mcp"}'Cursor, Windsurf and other
mcpServers-style clients (Zed, OpenCode and Goose use their own formats, see below):If a desktop app reports that
gemini-media-mcpwas not found, it doesn't see your shell's PATH: use the full path fromcommand -v gemini-media-mcp. Every client, plus Docker and the skills installer, is covered in packaging/INSTALL-SNIPPETS.md. -
If your agent doesn't forward environment variables (plugins often don't), store the key once:
To run it in Docker instead: docker run -i --rm --user "$(id -u):$(id -g)" -e GEMINI_API_KEY -v "$PWD/media:/output" -v gemini-media-state:/state ghcr.io/mordor-forge/gemini-media-mcp
(--user makes the generated files yours on Linux; the named /state volume keeps spend accounting and video jobs between runs).
Tools
Every result includes the saved file's path, a gemini-media://files/<name> URI and its cost. You can pass the URI as an input to another tool. Clients that can't read the server's disk (for example over HTTP) can fetch the file with resources/read.
Models
Use an alias or a full model ID. Run gemini-media-mcp models for the built-in catalog with prices, or call list_models, which also applies your override file and, with live: true, checks what your key can use.
Updating or adding a model without waiting for a release. Create a YAML file and point GEMINI_MEDIA_CATALOG at it. The server merges it with the built-in catalog by id and reloads it automatically:
See internal/catalog/models.yaml for the full schema.
Configuration
Settings are layered: built-in defaults < config file < environment < flags. The config file lives at ~/.config/gemini-media-mcp/config.yaml on Linux, ~/Library/Application Support/gemini-media-mcp/config.yaml on macOS, %AppData%\gemini-media-mcp\config.yaml on Windows, or wherever GEMINI_MEDIA_CONFIG points. Unknown keys produce warnings instead of errors, and get_config shows where each setting came from.
How the backend is chosen:
- An explicit
backendsetting wins. - Otherwise the SDK switches
GOOGLE_GENAI_USE_ENTERPRISEandGOOGLE_GENAI_USE_VERTEXAIdecide. - Otherwise an API key selects the Gemini API, even if
GOOGLE_CLOUD_PROJECTis set in your shell for other tools. - Otherwise a project selects Vertex AI.
On Vertex, an API key without a project uses express mode, which does not support video.
Spend and budgets
Google doesn't return costs, so the server computes them from its price table and the token usage the API reports. Every call records an estimate before it runs and the reconciled cost afterwards. Failed and safety-blocked generations count as $0. A response that used tokens but returned no media (for example MAX_TOKENS) is charged for those tokens.
- Results include
cost {estimatedUsd, usd, basis}.get_usageandgemini-media-mcp usagesummarize spend. The ledger is a plain JSONL file in the state directory. - Budgets are enforced before each call, across concurrent calls and across several server processes sharing a state directory (in-flight calls hold their estimate under a file lock).
- With
GEMINI_MEDIA_CONFIRM_ABOVE_USD=1, an expensive call (for example a $3.20 Veo clip) comes back as a[confirmation]error. The agent asks you, then retries withapprovedCostUsd. - These figures are estimates. For a hard stop, also set a spend cap in AI Studio or a budget in Cloud Billing.
HTTP mode
- HTTP mode is stateless Streamable HTTP, which MCP 2026-07-28 requires; older clients still work.
- The server protects against DNS rebinding and cross-origin requests.
- It refuses to listen on a non-loopback address without a token.
- It only reads input files from the output directory or
GEMINI_MEDIA_INPUT_DIRS(unlessGEMINI_MEDIA_ALLOW_ANY_INPUT_PATH=true). GET /healthzreports status.
Skills
The skills/ directory contains Agent Skills that teach agents the full workflow for each media type: intent, prompt craft, model choice, cost checks, review and iteration. They cover interactive use as well as unattended runs.
Plugin installs (Claude Code, Codex, Gemini CLI, VS Code) include the skills. To install them in any Agent Skills–compatible agent, run npx skills add mordor-forge/gemini-media-mcp, or copy the folders into .agents/skills/ or ~/.claude/skills/.
Development
Paid live E2E tests are run locally with your own API key. CI runs the free
checks and compiles/vets the E2E tests, but never executes live generation,
including after merges to main or on manual workflow runs.
- AGENTS.md: guide for coding agents and contributors.
- docs/architecture-review.md: architecture and design decisions.
- internal/catalog/models.yaml: model catalog. When Google changes models, edit this file, not the code.
License
來源:README.md,提交 747f3b8
工具
0版本歷史
1- v1.0.0最新Oct 8, 2026


