Google Cloud Storage Fuse

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

仅含说明

Cloud Storage FUSE (Advanced)

Cloud Storage FUSE (gcsfuse) is a POSIX file-system adapter over Cloud Storage's immutable object store. Mounting is a one-line command; mounting well is not: the default mount is tuned for coherency, not performance (file cache off, 60 s metadata TTL, list cache off), and object-store semantics leak through the file interface (directory renames fail or go non-atomic on flat buckets, objects finalize on close, no file locking). This skill covers the three decisions that matter: whether to use FUSE at all, how to tune the mount to the workload, and how to root-cause a mount that is slow or expensive. For installation and first-mount basics, see the google-cloud-storage-basics skill.

Attribution

Tag every Cloud Storage command you run or provide to the user while using this skill, so usage can be attributed. The tag identifies only the skill and its version; it carries no user data.

  • Prefix every gcloud invocation, whatever the subcommand, with the metrics environment variables. Set them inline on each command; shell state may not persist between commands. Use this append form verbatim. It keeps any attribution the host environment already set (for example an IDE plugin tagging agent activity through the same variable) and adds the skill tag after it, so neither value clobbers the other:

    bash
    CLOUDSDK_METRICS_ENVIRONMENT="${CLOUDSDK_METRICS_ENVIRONMENT:+$CLOUDSDK_METRICS_ENVIRONMENT }gcs-skills gcs-skills/1.0 (skill:google-cloud-storage-fuse)" \gcloud <command> [flags]

    Do not use gcloud config set for this: it would persist beyond the current task and mislabel unrelated usage.

  • On direct HTTP calls to the Cloud Storage APIs (for example with curl), set this exact User-Agent header, verbatim — the collection pipeline parses the gcs-skills/<version> and skill:<name> tokens, so any rewording breaks attribution:

    User-Agent: gcs-skills/1.0 (skill:google-cloud-storage-fuse)

Step 1 — Fit Gate (always run this first)

Never produce mount guidance before the fit gate. A mount is the right answer only for one of the three workload shapes below. If the workload's access pattern is unknown, ask — one question about whether the reading code can take gs:// paths usually settles it.

Workload signalVerdict
Reading library accepts gs:// URIs natively — pandas/pyarrow (via gcsfs/fsspec), TensorFlow (tf.io.gfile), or any fsspec/gcsfs-based loaderNative reads, no mount. Point the code at gs:// paths and stop.
Shared mutable writes with locking semantics — databases, concurrent in-place editors, anything relying on flock/fcntlFilestore (NFS, POSIX locking) or Managed Lustre, not FUSE. Stop.
Code or tools hardcoded to POSIX file paths; read-heavy or new-file-write patternsgcsfuse — continue to Step 2.

Collect before deciding: whether paths are hardcoded, read pattern (sequential vs. random, re-read frequency), write pattern (new files vs. edits vs. directory renames). These same signals drive tuning later — record the answers.

Step 2 — Route by intent

User intent (prompt shape)Go to
Provision: "mount my bucket for X", "get training data into my pods"GKE Training Deployment [blocked]
Safety/semantics: "is this write pattern safe?", "can multiple writers share the mount?"Checkpoint & Write Safety [blocked]
Regression: "training is slow", "the GCS bill spiked", "throughput dropped"Performance & Cost Diagnosis [blocked]

Never diagnose a regression without telemetry. If gcsfuse metrics are not enabled on the mount, enabling them is the first remediation step — the diagnosis reference starts there.

Reference Directory

  • GKE Training Deployment [blocked]: Fit-gated, performance-tuned mounts for training workloads — GKE CSI version gates, Workload Identity principal:// IAM bindings, profile StorageClasses vs. static PVs, file cache sizing on Local SSD, sidecar resource annotations, complete KSA/PVC/Job manifests, and the Compute Engine and Cloud Run variants.

  • Checkpoint & Write Safety [blocked]: Verdicts on write patterns — file vs. directory rename atomicity on flat vs. HNS buckets, close-vs-fsync finalization, concurrent-writer (ESTALE) semantics, streaming-write memory budgets, HNS migration, and the aiml-checkpointing profile.

  • Performance & Cost Diagnosis [blocked]: Telemetry-first runbook for slow mounts and bill spikes — enabling and reading gcsfuse metrics, mapping cache-hit and request-mix signatures to misconfigurations, the coherency-tuned defaults, tuned config keys with their staleness caveats, and billing-line (Class A/B) attribution.

来源与署名

来源:gemini-cli-extensions/data-agent-kit-starter-pack位于skills/google-cloud-storage-fuse提交2df10e2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

为 ETL、ELT 或数据集成任务规划源到目标的模式映射,产出文档化的映射宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修复区域级 Google Cloud MCP 服务器配置中未替换的区域占位符,使缺失的 MCP 工具得以注册。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分类215今天更新

Ml Best Practices

gemini-cli-extensions

为机器学习笔记本提供分步方案,涵盖聚类、预测、分类、回归和模型比较。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指导代理检测 Python 项目的依赖管理器并正确安装依赖,而不是使用全局 pip。

Software Development215今天更新

Google Cloud Auth Verification

gemini-cli-extensions

Mandatory Step 0 pre-flight execution order and authentication verification for Google Cloud Platform (GCP), Application Default Credentials (ADC), gcloud CLI, Spark, Dataproc, BigQuery, GCS, and notebook runtimes. Use whenever interacting with GCP resources, running Spark/PySpark pipelines, BigQuery queries, GCS paths (gs://), or creating/running notebooks.

待分类215今天更新