Gemma Multimodal Fine-Tuner
Skill by ara.so — Daily 2026 Skills collection.
Fine-tune Gemma 4 and Gemma 3n models on text, images, and audio data entirely on Apple Silicon (MPS), with support for streaming large datasets from GCS/BigQuery without filling local storage.
What It Does
- Text LoRA: instruction-tuning or completion fine-tuning from local CSV
- Image + Text LoRA: captioning and VQA from local CSV
- Audio + Text LoRA: the only Apple-Silicon-native path for this modality
- Cloud streaming: train on terabytes from GCS/BigQuery without local copy
- MPS-native: no NVIDIA GPU required — runs on MacBook Pro/Air/Mac Studio
Installation
Prerequisites
- macOS 12.3+ with Apple Silicon (arm64)
- Python 3.10+ (native arm64, not Rosetta)
- Hugging Face account with Gemma access
Authenticate with Hugging Face
CLI Commands
Configuration (config/config.ini)
The config is hierarchical INI: defaults → groups → models → datasets → profiles.
Use GEMMA_TUNER_CONFIG env var to point to config outside repo root:
Modality Configuration
Text-Only Fine-Tuning
Instruction tuning (user/assistant pairs):
Completion tuning (full sequence trained):
CSV format for instruction tuning (data/datasets/my-text-dataset/train.csv):
Image Fine-Tuning
CSV format (data/datasets/my-image-dataset/train.csv):
Audio Fine-Tuning
CSV format (data/datasets/my-audio-dataset/train.csv):
Supported Models
Add custom models with a [model:your-name] section using group = gemma.
Dataset Directory Layout
Output Layout
Python API Examples
Running Fine-Tuning Programmatically
Using Device Utilities
Loading and Inspecting Datasets
Custom LoRA Config
Common Patterns
Full Workflow: Text Instruction Tuning
GCS Streaming for Large Datasets
Set credentials:
Add a Custom Gemma Checkpoint
Troubleshooting
Wrong architecture (x86_64 instead of arm64)
MPS out of memory
- Reduce
batch_size(try 1) - Increase
gradient_accumulation_stepsto compensate - Use a smaller model (
e2binstead ofe4b) - Reduce
max_seq_length
Gemma 4 model not loading
Config not found outside repo root
Hugging Face auth errors
System check before debugging anything else
Audio tower loaded even for text-only runs
This is a known v1 issue — USM audio tower weights stay in memory even for modality = text. See README/KNOWN_ISSUES.md. Workaround: use a smaller model variant to stay within RAM budget.


