Gemma Gem Browser AI
Skill by ara.so — Daily 2026 Skills collection.
Gemma Gem is a Chrome extension that runs Google's Gemma 4 model entirely on-device via WebGPU. It injects a chat overlay into every page and exposes a tool-calling agent loop that can read pages, click elements, fill forms, execute JavaScript, and take screenshots — all without sending data to any server.
Architecture Overview
- Offscreen document (
offscreen/): Loads the ONNX model via@huggingface/transformers, runs the agent loop, streams tokens. - Service worker (
background/): Routes messages, handlestake_screenshotandrun_javascript. - Content script (
content/): Injects shadow DOM chat UI, executes DOM tools. agent/: Zero-dependency module definingModelBackendandToolExecutorinterfaces — extractable as a standalone library.
Install & Build
Load the extension:
- Open
chrome://extensions - Enable Developer mode
- Click Load unpacked → select
.output/chrome-mv3-dev/
Model download happens automatically on first chat open:
onnx-community/gemma-4-E2B-it-ONNX— ~500 MB (default)onnx-community/gemma-4-E4B-it-ONNX— ~1.5 GB
Models are cached in the browser's cache storage after the first run.
Key Interfaces (agent/)
ModelBackend
ToolExecutor
Agent Loop
Built-in Tools
Adding a New Tool
Tools live in two places: the definition (in the offscreen agent) and the executor (in content script or service worker).
Step 1 — Define the tool schema
Step 2 — Register in the tool list
Step 3 — Implement execution in the content script
Step 4 — Handle service-worker-side tools
Message Routing Pattern
The service worker acts as a message bus. All communication uses chrome.runtime.sendMessage.
Model Configuration
Settings & Persistence
Settings are stored via chrome.storage.sync:
Shadow DOM Chat UI Pattern
The content script injects a shadow DOM to isolate styles:
Debugging
All logs use [Gemma Gem] prefix. Development builds log info/debug/warn; production only logs errors.
Key things to check in offscreen document logs:
- Model download progress
- Full prompt construction
- Token counts per turn
- Raw model output (before tool call parsing)
- Tool execution results
Common Patterns & Gotchas
WebGPU availability check:
Offscreen document lifecycle — Chrome may suspend the offscreen document. Ping it before sending messages:
Context window management — Gemma 4 supports 128K tokens but inference slows with long contexts. Clear history per-page with clear_context or limit stored turns:
Tool call parsing — Gemma 4 emits tool calls in a structured format. If adding custom parsing, guard against partial/streamed JSON:
CSS selector safety for DOM tools:


