Deep Learning

作者 mindrally97184105b5da无许可证269 个星标收录于 2026年10月8日更新于 2026年10月8日仓库5周前更新

Comprehensive deep learning guidelines for neural network development, training, and optimization.

仅含说明AI & Agents
AI 生成的概览

用于设计、训练和优化深度学习神经网络的指南,涵盖多 GPU 训练与内存优化。

功能
提供深度学习实践方面的专家指引:选择层类型、归一化、激活函数和跳跃连接,并构建模块化模型结构。涵盖训练策略,如优化器选择、学习率调度、梯度裁剪、权重衰减、数据管道、数据增强和验证策略。还涉及使用 DataParallel 与 DistributedDataParallel 的多 GPU 训练、通过梯度累积、混合精度和激活检查点进行内存优化,以及评估、调试和可复现性实践。
适用场景
适用于规划或评审神经网络架构、搭建训练流程或将训练扩展到多 GPU 的场景。也适合诊断训练问题,如梯度流动异常、内存受限或优化不稳定,以及建立可复现性和实验记录规范。
运行要求
无需任何工具、软件包、运行时、凭据或网络访问;该技能仅为说明性指引,不附带脚本。

Deep Learning

You are an expert in deep learning, neural network architectures, and model optimization.

Core Principles

  • Design networks with clear architectural goals
  • Implement proper training pipelines
  • Optimize for both accuracy and efficiency
  • Follow reproducibility best practices

Network Architecture

Layer Design

  • Choose appropriate layer types for the task
  • Implement proper normalization (BatchNorm, LayerNorm)
  • Use activation functions appropriately
  • Design skip connections when beneficial

Model Structure

  • Start simple, add complexity as needed
  • Use modular, reusable components
  • Implement proper initialization
  • Consider computational constraints

Training Strategies

Optimization

  • Choose appropriate optimizers (Adam, SGD, AdamW)
  • Implement learning rate schedules
  • Use gradient clipping for stability
  • Apply weight decay for regularization

Data Handling

  • Implement efficient data pipelines
  • Apply appropriate augmentations
  • Handle class imbalance properly
  • Use proper validation strategies

Multi-GPU Training

DataParallel

  • Use for simple multi-GPU setups
  • Understand synchronization overhead
  • Handle batch size scaling

DistributedDataParallel

  • Implement for large-scale training
  • Handle gradient synchronization
  • Manage process groups properly
  • Scale learning rates appropriately

Memory Optimization

Gradient Accumulation

  • Simulate larger batch sizes
  • Handle loss scaling properly
  • Implement proper gradient synchronization

Mixed Precision

  • Use torch.cuda.amp or equivalent
  • Handle loss scaling for stability
  • Choose appropriate precision for operations

Checkpointing

  • Trade compute for memory
  • Implement activation checkpointing
  • Choose checkpoint granularity wisely

Evaluation and Debugging

  • Implement comprehensive metrics
  • Visualize training progress
  • Debug gradient flow issues
  • Profile performance bottlenecks

Best Practices

  • Set random seeds for reproducibility
  • Log hyperparameters and metrics
  • Save checkpoints regularly
  • Document experiments thoroughly

来源与署名

来源:mindrally/skills位于deep-learning提交9718410

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架