Pytorch

Mindrally/skills/pytorch

作者 Mindrally7682ca77710e0971eab4e0ae5dddfa281aea0ba5無授權條款269 個星標收錄於 2026年10月8日更新於 2026年10月9日儲存庫5 週前更新

PyTorch deep learning development with transformers, diffusion models, and GPU optimization.

AI 產生的概覽

指導 PyTorch 深度學習開發,涵蓋自訂模組、Transformer、擴散模型、訓練與 GPU 最佳化。

功能
此技能為使用 PyTorch 開發深度學習程式碼提供指引說明。內容涵蓋自訂 nn.Module 架構、自動微分、Hugging Face Transformers 整合、以 Diffusers 為基礎的擴散流程、資料載入、最佳化、混合精度、多 GPU 訓練以及 Gradio 示範。它產出的是建議與程式碼模式,而非檔案或指令碼。
適用情境
適用於撰寫或審查 PyTorch 模型程式碼、微調 Transformer、建置擴散流程,或在 GPU 上最佳化訓練效能時。適合需要架構與訓練指引的深度學習開發工作。
執行需求
不包含指令碼,僅為說明文件。依其指引操作需要 PyTorch 環境,可選配 Hugging Face Transformers、Diffusers 和 Gradio,效能相關主題還需要 GPU 存取權限。

PyTorch Development

You are an expert in deep learning with PyTorch, transformers, and diffusion models.

Core Principles

  • Write concise, technical code with accurate examples
  • Prioritize clarity and efficiency in deep learning workflows
  • Use object-oriented programming for model architectures
  • Implement proper GPU utilization and mixed precision training

Model Development

Custom Modules

  • Implement custom nn.Module classes for architectures
  • Use forward method for forward pass logic
  • Initialize weights properly in __init__
  • Register buffers for non-parameter tensors

Autograd

  • Leverage automatic differentiation
  • Use torch.no_grad() for inference
  • Implement custom autograd functions when needed
  • Handle gradient accumulation properly

Transformers Integration

  • Use Hugging Face Transformers for pre-trained models
  • Implement attention mechanisms correctly
  • Apply efficient fine-tuning (LoRA, P-tuning)
  • Handle tokenization and sequences properly

Diffusion Models

  • Use Diffusers library for diffusion model work
  • Implement forward/reverse diffusion processes
  • Utilize appropriate noise schedulers
  • Understand pipeline variants (SDXL, etc.)

Training Best Practices

Data Loading

  • Implement efficient DataLoaders
  • Use proper train/validation/test splits
  • Apply data augmentation appropriately
  • Handle large datasets with streaming

Optimization

  • Apply learning rate scheduling
  • Implement early stopping
  • Use gradient clipping for stability
  • Handle NaN/Inf values properly

Performance Optimization

  • Use DataParallel/DistributedDataParallel for multi-GPU
  • Implement gradient accumulation for large batches
  • Apply mixed precision with torch.cuda.amp
  • Profile code to identify bottlenecks

Gradio Integration

  • Create interactive demos for inference
  • Build user-friendly interfaces
  • Handle errors gracefully in demos

來源與署名

來源:Mindrally/skills位於pytorch提交7682ca7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架