Gemini Media

io.github.mordor-forgev1.0.0Updated Oct 8, 2026

Generate images (Nano Banana), video (Veo, Gemini Omni), speech and music (Lyria) with Google models

VerifiedSTDIODesktop onlyAI & MLMedia & Design

Overview

AI-generated overview

Lets an assistant generate and edit images, video, speech and music with Google's generative media models, with cost estimates and budgets.

What it does
Wraps Google's generative media models for image generation and editing, video generation, extension and editing, text-to-speech and music creation. Tools include generate_image, edit_image, generate_video, get_video, extend_video, edit_video, generate_speech, generate_music, list_models, estimate_cost, get_usage and get_config. Results report the saved file path, a gemini-media://files/ URI that can be chained into other tools, and an estimated cost.
When to use it
Useful when an assistant needs to produce or modify media assets such as images, clips, voiceovers or music, including multi-asset production workflows. Also suited to cost-sensitive use, since it estimates spend before running a request and enforces session, daily and monthly budgets.
Requirements
A local binary on PATH, or the Docker image, Claude Desktop bundle or Gemini CLI extension. Credentials are either a Gemini API key (GEMINI_API_KEY) or a Google Cloud project with Vertex AI enabled (GOOGLE_CLOUD_PROJECT) plus Application Default Credentials. Network access to Google APIs is required; media is written to an output directory.
Before you install
Generation calls are billed by Google; the server's cost figures are estimates, so a hard cap should also be set in AI Studio or Cloud Billing. Credentials such as GEMINI_API_KEY are stored in config when using the configure command. Tools write media files to disk, and HTTP mode can read input files from configured directories.

Installation

In SourceWeft

  1. Open Gemini Media in the dashboard and add it to a workspace.
  2. Enable the server for the chats that should use its tools.

Desktop only via STDIO. STDIO servers start a local process, so they need the SourceWeft desktop host.

Other MCP clients

Follow the launch instructions in the repository.

README

gemini-media-mcp

[CI] [Go] [License]

An MCP server for Google's generative media models: images (Nano Banana 2.1 / Pro), video (Veo 3.1 and Gemini Omni Flash), speech (Gemini 3.8 TTS) and music (Lyria 3.5). It ships as a single Go binary, speaks stdio and Streamable HTTP (MCP 2026-07-28), works with the Gemini API or Vertex AI, and comes with agent skills and plugin packaging for Claude Code, Codex, Gemini CLI, VS Code/Copilot, Cursor and more.

  • Current models, updated without a release. A built-in catalog records IDs, aliases, lifecycle, parameters and prices. Retired models redirect to their replacement, and new model IDs work before the catalog knows them. You can override or extend the catalog with a hot-reloaded YAML file.
  • Cost-aware. Every result reports its estimated cost, and estimate_cost compares options before you spend. Spend is recorded in a ledger, capped by session, daily and monthly budgets, and calls above a threshold need explicit approval.
  • Agent-friendly. Each tool matches a workflow, and errors come back as [kind] message + Hint. Image results include inline previews, and outputs can be chained by URI. Video runs as async jobs with long-polling and progress notifications.
  • Robust. Backend and credentials are detected the way the Google SDKs do it. Vertex locations are chosen per model, retries and timeouts are built in, files are written atomically with provenance, and HTTP mode ships with security defaults.

Upgrading from v0? The tools changed. See the migration table. The review also covers what was broken, why, and the design of v1.

Quick start

  1. Get credentials. Either:

    • an API key from Google AI Studio (GEMINI_API_KEY), or
    • a Google Cloud project with Vertex AI enabled (GOOGLE_CLOUD_PROJECT plus gcloud auth application-default login).
  2. Install the binary. On macOS or Linux:

    bash
    mkdir -p ~/.local/bin   # must be on your PATHos=$(uname -s | tr '[:upper:]' '[:lower:]'); arch=$(uname -m | sed 's/x86_64/x64/; s/aarch64/arm64/')curl -fsSL "https://github.com/mordor-forge/gemini-media-mcp/releases/latest/download/$os.$arch.gemini-media-mcp.tar.gz" \  | tar -xz -C ~/.local/bin gemini-media-mcpgemini-media-mcp version

    On Windows, download the windows zip from the latest release and put gemini-media-mcp.exe on your PATH. With Go 1.26+: go install github.com/mordor-forge/gemini-media-mcp/cmd/gemini-media-mcp@latest.

    The Gemini CLI extension, the Claude Desktop bundle and the Docker image include the binary, so they skip this step.

  3. Add the server to your agent. Every config below runs gemini-media-mcp from your PATH.

    Claude Code (plugin: server plus skills):

    /plugin marketplace add mordor-forge/gemini-media-mcp/plugin install gemini-media@mordor-forge

    or just the server: claude mcp add gemini-media -e GEMINI_API_KEY=... -- gemini-media-mcp

    Codex (~/.codex/config.toml):

    toml
    [mcp_servers.gemini-media]command = "gemini-media-mcp"env_vars = ["GEMINI_API_KEY", "GOOGLE_API_KEY", "GOOGLE_CLOUD_PROJECT", "GOOGLE_CLOUD_LOCATION"]tool_timeout_sec = 300   # 4K images and video polls can exceed the 60 s default

    Gemini CLI (paid API keys, Vertex AI and Gemini Code Assist only, since June 2026): gemini extensions install https://github.com/mordor-forge/gemini-media-mcp

    Antigravity CLI: add the server to ~/.gemini/config/mcp_config.json, as shown in the install snippets.

    Claude Desktop: download gemini-media-mcp-<version>.mcpb from the latest release and open it.

    VS Code / Copilot: code --add-mcp '{"name":"gemini-media","command":"gemini-media-mcp"}'

    Cursor, Windsurf and other mcpServers-style clients (Zed, OpenCode and Goose use their own formats, see below):

    json
    { "mcpServers": { "gemini-media": { "command": "gemini-media-mcp", "env": { "GEMINI_API_KEY": "..." } } } }

    If a desktop app reports that gemini-media-mcp was not found, it doesn't see your shell's PATH: use the full path from command -v gemini-media-mcp. Every client, plus Docker and the skills installer, is covered in packaging/INSTALL-SNIPPETS.md.

  4. If your agent doesn't forward environment variables (plugins often don't), store the key once:

    bash
    echo "$GEMINI_API_KEY" | gemini-media-mcp configure --api-key-stdingemini-media-mcp doctor   # checks credentials, backend and model availability

To run it in Docker instead: docker run -i --rm --user "$(id -u):$(id -g)" -e GEMINI_API_KEY -v "$PWD/media:/output" -v gemini-media-state:/state ghcr.io/mordor-forge/gemini-media-mcp (--user makes the generated files yours on Linux; the named /state volume keeps spend accounting and video jobs between runs).

Tools

ToolWhat it does
generate_imageText-to-image with up to 14 reference images, 1K–4K, many aspect ratios, 1–4 variations, optional Google Search grounding
edit_imageChange an existing image (add/remove/restyle/relight/outpaint) while keeping the rest
generate_videoClip with native audio from text, a first frame, first+last frames, or reference images: Veo (4–8 s, up to 3 references) or Gemini Omni Flash (omni: 3–10 s, 360p drafts to 4K, up to 10 references). Returns a jobId
get_videoWait for a job (long-poll, default 45 s); downloads the video when done. Safe to repeat
extend_videoContinue a finished clip: Omni adds up to 10 s (40 s total), Veo about 7 s (up to 148 s total)
edit_videoChange a finished clip or a video file of up to 10 s with an instruction (Gemini Omni)
generate_speechText-to-speech (WAV): one voice, or a two-speaker dialogue, with per-line style control and 30 voices
generate_music30-second clips or full songs with lyrics, structure tags, tempo, instrumental mode and image inspiration
list_modelsCurrent models, aliases, status, prices; detail for supported parameters, live to check what your key can use
estimate_costPrice a request before running it and compare models
get_usageEstimated spend by period, model and tool; budgets remaining; running jobs
get_configActive backend and why it was chosen, output directory, defaults, warnings

Every result includes the saved file's path, a gemini-media://files/<name> URI and its cost. You can pass the URI as an input to another tool. Clients that can't read the server's disk (for example over HTTP) can fetch the file with resources/read.

Models

Use an alias or a full model ID. Run gemini-media-mcp models for the built-in catalog with prices, or call list_models, which also applies your override file and, with live: true, checks what your key can use.

MediaAliases (default first)
Imagenb2 (Nano Banana 2.1, default), pro (Nano Banana Pro: highest fidelity), nb2-lite (cheapest inputs)
Videolite (cheapest Veo), omni (Gemini Omni Flash: prompt adherence, editing, extension to 40 s; Gemini API only), fast (Veo 4K, references, extension), standard (highest Veo quality)
Speechtts (Gemini 3.8 Flash TTS), tts-lite, tts-2.5, tts-pro
Musicclip (30 s), full (Lyria 3.5 songs)

Updating or adding a model without waiting for a release. Create a YAML file and point GEMINI_MEDIA_CATALOG at it. The server merges it with the built-in catalog by id and reloads it automatically:

yaml
defaults:  video: fastmodels:  - id: veo-3.1-fast-generate-preview    pricing: { perSecond: { 720p: 0.10, 1080p: 0.12, 4k: 0.30 } }  - id: veo-4.0-generate-preview        # a brand-new model    aliases: [veo4]    pricing: { perSecond: { 720p: 0.50 } }

See internal/catalog/models.yaml for the full schema.

Configuration

Settings are layered: built-in defaults < config file < environment < flags. The config file lives at ~/.config/gemini-media-mcp/config.yaml on Linux, ~/Library/Application Support/gemini-media-mcp/config.yaml on macOS, %AppData%\gemini-media-mcp\config.yaml on Windows, or wherever GEMINI_MEDIA_CONFIG points. Unknown keys produce warnings instead of errors, and get_config shows where each setting came from.

Environment variableConfig keyDefaultPurpose
GEMINI_API_KEY / GOOGLE_API_KEY / GEMINI_MEDIA_API_KEYapiKey–Gemini API key (or Vertex express-mode key)
GEMINI_MEDIA_PROJECT / GOOGLE_CLOUD_PROJECTproject–Vertex AI project (Application Default Credentials)
GEMINI_MEDIA_LOCATION / GOOGLE_CLOUD_LOCATION / GOOGLE_CLOUD_REGIONlocationper modelVertex region; models not offered there use their catalog location
GOOGLE_GENAI_USE_VERTEXAI / GOOGLE_GENAI_USE_ENTERPRISE––true selects Vertex AI, false the Gemini API (same semantics as the Google SDKs)
GEMINI_MEDIA_BACKENDbackendautoauto, gemini-api or vertex
MEDIA_OUTPUT_DIR / GEMINI_MEDIA_OUTPUT_DIRoutputDir~/generated_mediaWhere media is saved
GEMINI_MEDIA_STATE_DIRstateDir$XDG_STATE_HOME/gemini-media-mcp, else ~/.local/state/gemini-media-mcp (Linux) or the config directorySpend ledger and video jobs
GEMINI_MEDIA_BUDGET_SESSION_USD / _DAILY_USD / _MONTHLY_USDbudget.*UsdnoneSpend caps (estimated)
GEMINI_MEDIA_CONFIRM_ABOVE_USDbudget.confirmAboveUsdnoneCalls above this need approvedCostUsd
GEMINI_MEDIA_IMAGE_MODEL / _VIDEO_MODEL / _SPEECH_MODEL / _MUSIC_MODEL / GEMINI_MEDIA_VOICEdefaults.*catalogDefault models and voice
GEMINI_MEDIA_CATALOGcatalogFile–Catalog override file (hot-reloaded)
GEMINI_MEDIA_INPUT_DIRSinputDirs–Extra directories inputs may be read from (HTTP mode)
GEMINI_MEDIA_ALLOW_ANY_INPUT_PATHallowAnyInputPathtrue on stdio, false on HTTPRead input files from anywhere on disk
GEMINI_MEDIA_INLINE_PREVIEWS / GEMINI_MEDIA_PREVIEW_MAX_PIXELSinlinePreviews / previewMaxPixelstrue / 768Attach a downscaled preview to image results, and its longest side
GEMINI_MEDIA_REQUEST_TIMEOUT_SECONDSrequestTimeoutSeconds600Timeout of each Google API request
GEMINI_MEDIA_MAX_VIDEO_WAIT_SECONDSmaxVideoWaitSeconds600Upper limit for waitSeconds on video tools
GEMINI_MEDIA_RETRY_ATTEMPTSretry.attempts4Attempts for rate-limited (429) and unavailable (503) responses
GEMINI_MEDIA_LOG_LEVELlogLevelinfodebug, info, warn or error (logs go to stderr)
GEMINI_MEDIA_TRANSPORTtransportstdiostdio or http
GEMINI_MEDIA_HTTP_ADDR / GEMINI_MEDIA_HTTP_TOKENhttp.addr / http.authToken127.0.0.1:8765 / –HTTP listen address and bearer token
–http.path/mcpHTTP endpoint path
GEMINI_MEDIA_HTTP_ALLOWED_ORIGINShttp.allowedOrigins–Extra trusted browser origins, comma-separated
GEMINI_MEDIA_HTTP_STATEFULhttp.statefulfalseKeep per-client sessions (only for legacy 2025-11-25 clients that need them)

How the backend is chosen:

  1. An explicit backend setting wins.
  2. Otherwise the SDK switches GOOGLE_GENAI_USE_ENTERPRISE and GOOGLE_GENAI_USE_VERTEXAI decide.
  3. Otherwise an API key selects the Gemini API, even if GOOGLE_CLOUD_PROJECT is set in your shell for other tools.
  4. Otherwise a project selects Vertex AI.

On Vertex, an API key without a project uses express mode, which does not support video.

Spend and budgets

Google doesn't return costs, so the server computes them from its price table and the token usage the API reports. Every call records an estimate before it runs and the reconciled cost afterwards. Failed and safety-blocked generations count as $0. A response that used tokens but returned no media (for example MAX_TOKENS) is charged for those tokens.

  • Results include cost {estimatedUsd, usd, basis}. get_usage and gemini-media-mcp usage summarize spend. The ledger is a plain JSONL file in the state directory.
  • Budgets are enforced before each call, across concurrent calls and across several server processes sharing a state directory (in-flight calls hold their estimate under a file lock).
  • With GEMINI_MEDIA_CONFIRM_ABOVE_USD=1, an expensive call (for example a $3.20 Veo clip) comes back as a [confirmation] error. The agent asks you, then retries with approvedCostUsd.
  • These figures are estimates. For a hard stop, also set a spend cap in AI Studio or a budget in Cloud Billing.

HTTP mode

bash
GEMINI_API_KEY=... gemini-media-mcp serve --transport http                 # http://127.0.0.1:8765/mcpGEMINI_MEDIA_HTTP_TOKEN=secret gemini-media-mcp serve --transport http --http-addr 0.0.0.0:8765claude mcp add --transport http gemini-media http://127.0.0.1:8765/mcp --header "Authorization: Bearer secret"
  • HTTP mode is stateless Streamable HTTP, which MCP 2026-07-28 requires; older clients still work.
  • The server protects against DNS rebinding and cross-origin requests.
  • It refuses to listen on a non-loopback address without a token.
  • It only reads input files from the output directory or GEMINI_MEDIA_INPUT_DIRS (unless GEMINI_MEDIA_ALLOW_ANY_INPUT_PATH=true).
  • GET /healthz reports status.

Skills

The skills/ directory contains Agent Skills that teach agents the full workflow for each media type: intent, prompt craft, model choice, cost checks, review and iteration. They cover interactive use as well as unattended runs.

SkillFor
gemini-imageImages: generation, editing, multi-reference composition, text rendering
gemini-videoVideo: text/image-to-video, frame interpolation, reference ingredients, Omni editing, extension, async jobs
gemini-speechVoiceovers, narration, two-speaker dialogue, voice and style selection
gemini-musicClips and full songs with structure, lyrics and tempo
gemini-media-productionMulti-asset projects (storyboard → keyframes → video → voiceover → music → ffmpeg assembly) with a budget plan

Plugin installs (Claude Code, Codex, Gemini CLI, VS Code) include the skills. To install them in any Agent Skills–compatible agent, run npx skills add mordor-forge/gemini-media-mcp, or copy the folders into .agents/skills/ or ~/.claude/skills/.

Development

Paid live E2E tests are run locally with your own API key. CI runs the free checks and compiles/vets the E2E tests, but never executes live generation, including after merges to main or on manual workflow runs.

bash
go build ./...go test -race ./...golangci-lint runpython3 scripts/validate-skills.py skillsGEMINI_MEDIA_E2E=1 GEMINI_API_KEY=... go test -tags=e2e ./internal/media/ -run E2E -v -timeout 30m   # live, costs cents# add GEMINI_MEDIA_E2E_VIDEO=1 for the video test (~$0.20), GEMINI_MEDIA_E2E_OMNI=1 for the Omni clip + edit (~$0.30),# GEMINI_MEDIA_E2E_OUTPUT_DIR=./e2e-out to keep the files

License

Apache-2.0

Source: README.md at commit 747f3b8

Tools

0
Tool metadata has not been indexed yet.

Version history

1
  1. v1.0.0LatestOct 8, 2026