Gemma Tuner Multimodal

作者 reason-machines2384a003145a无许可证83 个星标收录于 2026年10月8日更新于 2026年10月8日仓库3个月前更新

Fine-tune Gemma 4 and 3n models with audio, images, and text on Apple Silicon using PyTorch and Metal Performance Shaders.

仅含说明AI & Agents
AI 生成的概览

在 Apple Silicon 上使用 LoRA 和 MPS 对 Gemma 3n 与 Gemma 4 模型进行文本、图像或音频微调。

功能
指导在 Apple Silicon 上通过 PyTorch 和 Metal Performance Shaders 微调 Gemma 3n 与 Gemma 4 模型,涵盖文本指令或补全微调、图像描述与视觉问答,以及音频转写。它通过分层 INI 文件配置 LoRA 适配器,支持从 GCS 或 BigQuery 流式读取大型数据集,并生成包含检查点、指标和可导出的合并 Hugging Face SafeTensors 权重的训练运行。文档还介绍了用于系统检查、数据集准备、评估和故障排查的命令行命令。
适用场景
当你希望在 Mac 上本地微调 Gemma 多模态模型且没有 NVIDIA GPU 时使用。它适用于基于本地 CSV 文本、图像或音频数据集的 LoRA 训练,以及从 GCS 或 BigQuery 流式读取的大型云端数据集。
运行要求
macOS 12.3+ 的 Apple Silicon(arm64)、以原生方式安装的 Python 3.10+、PyTorch 与 torchaudio、gemma-tuner-multimodal 软件包,以及具有 Gemma 访问权限的 Hugging Face 账号(HF_TOKEN 或 huggingface-cli login)。Gemma 4 支持需要额外的依赖文件并建议使用独立虚拟环境。GCS/BigQuery 流式读取需要 Google 应用凭据和网络访问。该技能仅为说明文档,不附带脚本。

Gemma Multimodal Fine-Tuner

Skill by ara.so — Daily 2026 Skills collection.

Fine-tune Gemma 4 and Gemma 3n models on text, images, and audio data entirely on Apple Silicon (MPS), with support for streaming large datasets from GCS/BigQuery without filling local storage.


What It Does

  • Text LoRA: instruction-tuning or completion fine-tuning from local CSV
  • Image + Text LoRA: captioning and VQA from local CSV
  • Audio + Text LoRA: the only Apple-Silicon-native path for this modality
  • Cloud streaming: train on terabytes from GCS/BigQuery without local copy
  • MPS-native: no NVIDIA GPU required — runs on MacBook Pro/Air/Mac Studio

Installation

Prerequisites

  • macOS 12.3+ with Apple Silicon (arm64)
  • Python 3.10+ (native arm64, not Rosetta)
  • Hugging Face account with Gemma access
bash
# Install Python 3.12 if neededbrew install [email protected]
# Create venvpython3.12 -m venv .venvsource .venv/bin/activate
# Verify arm64 (must show arm64, not x86_64)python -c "import platform; print(platform.machine())"
# Install PyTorchpip install torch torchaudio
# Clone and installgit clone https://github.com/mattmireles/gemma-tuner-multimodalcd gemma-tuner-multimodalpip install -e .
# For Gemma 4 support (separate venv recommended)pip install -r requirements/requirements-gemma4.txt

Authenticate with Hugging Face

bash
huggingface-cli login# Or set environment variable:export HF_TOKEN=your_token_here

CLI Commands

bash
# Check system is readygemma-macos-tuner system-check
# Guided setup wizard (recommended for first run)gemma-macos-tuner wizard
# Prepare datasetgemma-macos-tuner prepare <dataset-profile>
# Fine-tune a modelgemma-macos-tuner finetune <profile> --json-logging
# Evaluate a rungemma-macos-tuner evaluate <profile-or-run>
# Export merged HF/SafeTensors (merges LoRA when adapter_config.json present)gemma-macos-tuner export <run-dir-or-profile>
# Blacklist bad samples from errorsgemma-macos-tuner blacklist <profile>
# List training runsgemma-macos-tuner runs list

Configuration (config/config.ini)

The config is hierarchical INI: defaults → groups → models → datasets → profiles.

ini
[defaults]output_dir = outputbatch_size = 2gradient_accumulation_steps = 8learning_rate = 2e-4num_train_epochs = 3
[model:gemma-3n-e2b-it]group = gemmabase_model = google/gemma-3n-E2B-it
[model:gemma-4-e2b-it]group = gemmabase_model = google/gemma-4-E2B-it
[dataset:my-audio-dataset]data_dir = data/datasets/my-audio-datasetaudio_column = audio_pathtext_column = transcript
[profile:my-audio-profile]model = gemma-3n-e2b-itdataset = my-audio-datasetmodality = audiolora_r = 16lora_alpha = 32lora_dropout = 0.05max_seq_length = 512

Use GEMMA_TUNER_CONFIG env var to point to config outside repo root:

bash
export GEMMA_TUNER_CONFIG=/path/to/my/config.ini

Modality Configuration

Text-Only Fine-Tuning

Instruction tuning (user/assistant pairs):

ini
[profile:text-instruction]model = gemma-3n-e2b-itdataset = my-text-datasetmodality = texttext_sub_mode = instructionprompt_column = prompttext_column = responsemax_seq_length = 2048lora_r = 16lora_alpha = 32

Completion tuning (full sequence trained):

ini
[profile:text-completion]model = gemma-3n-e2b-itdataset = my-text-datasetmodality = texttext_sub_mode = completiontext_column = textmax_seq_length = 2048

CSV format for instruction tuning (data/datasets/my-text-dataset/train.csv):

csv
prompt,response"What is photosynthesis?","Photosynthesis is the process by which plants...""Explain LoRA fine-tuning","LoRA (Low-Rank Adaptation) is a parameter-efficient..."

Image Fine-Tuning

ini
[profile:image-caption]model = gemma-3n-e2b-itdataset = my-image-datasetmodality = imageimage_sub_mode = captioningimage_token_budget = 256prompt_column = prompttext_column = captionmax_seq_length = 512

CSV format (data/datasets/my-image-dataset/train.csv):

csv
image_path,prompt,caption/data/images/img1.jpg,Describe this image,A dog sitting on a green lawn.../data/images/img2.jpg,What is shown here,A bar chart showing quarterly revenue...

Audio Fine-Tuning

ini
[profile:audio-asr]model = gemma-3n-e2b-itdataset = my-audio-datasetmodality = audioaudio_column = audio_pathtext_column = transcriptmax_seq_length = 512lora_r = 16lora_alpha = 32lora_dropout = 0.05

CSV format (data/datasets/my-audio-dataset/train.csv):

csv
audio_path,transcript/data/audio/recording1.wav,The patient presents with acute respiratory symptoms/data/audio/recording2.wav,Counsel objects to the characterization of the evidence

Supported Models

Model KeyHugging Face IDNotes
gemma-3n-e2b-itgoogle/gemma-3n-E2B-itDefault, ~2B instruct
gemma-3n-e4b-itgoogle/gemma-3n-E4B-it~4B instruct
gemma-4-e2b-itgoogle/gemma-4-E2B-itNeeds requirements-gemma4.txt
gemma-4-e4b-itgoogle/gemma-4-E4B-itNeeds requirements-gemma4.txt
gemma-4-e2bgoogle/gemma-4-E2BBase, needs Gemma 4 stack
gemma-4-e4bgoogle/gemma-4-E4BBase, needs Gemma 4 stack

Add custom models with a [model:your-name] section using group = gemma.


Dataset Directory Layout

data/└── datasets/    └── <dataset-name>/        ├── train.csv       # required        ├── validation.csv  # optional        └── test.csv        # optional

Output Layout

output/└── {run-id}-{profile}/    ├── metadata.json    ├── metrics.json    ├── checkpoint-*/    └── adapter_model/      # LoRA artifacts

Python API Examples

Running Fine-Tuning Programmatically

python
from gemma_tuner.core.config import load_configfrom gemma_tuner.core.ops import run_finetune
# Load configconfig = load_config("config/config.ini")
# Run fine-tuning for a profilerun_finetune(profile="my-audio-profile", config=config, json_logging=True)

Using Device Utilities

python
from gemma_tuner.utils.device import get_device, memory_hint
device = get_device()   # Returns "mps", "cuda", or "cpu"print(f"Training on: {device}")
hint = memory_hint(model_key="gemma-3n-e2b-it")print(hint)

Loading and Inspecting Datasets

python
from gemma_tuner.utils.dataset_utils import load_csv_dataset
train_df, val_df = load_csv_dataset(    data_dir="data/datasets/my-text-dataset",    text_column="response",    prompt_column="prompt")print(f"Train samples: {len(train_df)}, Val samples: {len(val_df)}")

Custom LoRA Config

python
from peft import LoraConfig, get_peft_modelfrom transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(    "google/gemma-3n-E2B-it",    torch_dtype="auto",    device_map="mps")
lora_config = LoraConfig(    r=16,    lora_alpha=32,    lora_dropout=0.05,    target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],    task_type="CAUSAL_LM")
model = get_peft_model(model, lora_config)model.print_trainable_parameters()

Common Patterns

Full Workflow: Text Instruction Tuning

bash
# 1. Prepare your datamkdir -p data/datasets/my-datasetcp train.csv data/datasets/my-dataset/cp validation.csv data/datasets/my-dataset/
# 2. Add profile to config/config.inicat >> config/config.ini << 'EOF'[dataset:my-dataset]data_dir = data/datasets/my-dataset
[profile:my-text-run]model = gemma-3n-e2b-itdataset = my-datasetmodality = texttext_sub_mode = instructionprompt_column = prompttext_column = responsemax_seq_length = 2048lora_r = 16lora_alpha = 32EOF
# 3. Prepare datasetgemma-macos-tuner prepare my-dataset
# 4. Fine-tunegemma-macos-tuner finetune my-text-run --json-logging
# 5. Export merged weightsgemma-macos-tuner export my-text-run

GCS Streaming for Large Datasets

ini
[dataset:large-audio-gcs]source = gcsgcs_bucket = my-bucketgcs_prefix = audio-training-data/audio_column = audio_pathtext_column = transcript
[profile:large-audio-run]model = gemma-3n-e4b-itdataset = large-audio-gcsmodality = audiolora_r = 32lora_alpha = 64

Set credentials:

bash
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account.jsongemma-macos-tuner finetune large-audio-run

Add a Custom Gemma Checkpoint

ini
[model:my-custom-gemma]group = gemmabase_model = my-org/my-gemma-checkpoint
[profile:custom-run]model = my-custom-gemmadataset = my-datasetmodality = texttext_sub_mode = instruction

Troubleshooting

Wrong architecture (x86_64 instead of arm64)

bash
python -c "import platform; print(platform.machine())"# Must be arm64 — if x86_64, reinstall Python natively:brew install [email protected]python3.12 -m venv .venv && source .venv/bin/activate

MPS out of memory

  • Reduce batch_size (try 1)
  • Increase gradient_accumulation_steps to compensate
  • Use a smaller model (e2b instead of e4b)
  • Reduce max_seq_length

Gemma 4 model not loading

bash
# Gemma 4 requires the updated Transformers stackpip install -r requirements/requirements-gemma4.txt# Use a separate venv if you also need Gemma 3n

Config not found outside repo root

bash
export GEMMA_TUNER_CONFIG=/absolute/path/to/config/config.inigemma-macos-tuner finetune my-profile

Hugging Face auth errors

bash
huggingface-cli login# Or:export HF_TOKEN=your_hf_token# Accept Gemma license at: https://huggingface.co/google/gemma-3n-E2B-it

System check before debugging anything else

bash
gemma-macos-tuner system-check

Audio tower loaded even for text-only runs

This is a known v1 issue — USM audio tower weights stay in memory even for modality = text. See README/KNOWN_ISSUES.md. Workaround: use a smaller model variant to stay within RAM budget.


Architecture Reference

FileRole
gemma_tuner/cli_typer.pyMain CLI entrypoint (gemma-macos-tuner)
gemma_tuner/core/ops.pyDispatches prepare/finetune/evaluate/export
gemma_tuner/scripts/finetune.pyRouter: Gemma models → models/gemma/finetune.py
gemma_tuner/models/gemma/finetune.pyCore training loop with LoRA
gemma_tuner/scripts/export.pyMerges LoRA → HF/SafeTensors tree
gemma_tuner/utils/device.pyMPS/CUDA/CPU selection and memory hints
gemma_tuner/utils/dataset_utils.pyCSV loading, blacklist/protection semantics
gemma_tuner/wizard/Interactive CLI wizard (questionary + Rich)
config/config.iniHierarchical INI configuration

来源与署名

来源:reason-machines/trending-skills位于skills/gemma-tuner-multimodal提交2384a00

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架