Customize Model Deployment
Interactive guided workflow for deploying Azure OpenAI models with full customization control over version, SKU, capacity, content filtering, and advanced options.
Quick Reference
When to Use This Skill
Use this skill when you need precise control over deployment configuration:
- ✅ Choose specific model version (not just latest)
- ✅ Select deployment SKU (GlobalStandard vs Standard vs PTU)
- ✅ Set exact capacity within available range
- ✅ Configure content filtering (RAI policy selection)
- ✅ Enable advanced features (dynamic quota, priority processing, spillover)
- ✅ PTU deployments (Provisioned Throughput Units)
Alternative: Use preset for quick deployment to the best available region with automatic configuration.
Comparison: customize vs preset
Prerequisites
- Azure subscription with Cognitive Services Contributor or Owner role
- Microsoft Foundry project resource ID (format:
/subscriptions/{sub}/resourceGroups/{rg}/providers/Microsoft.CognitiveServices/accounts/{account}/projects/{project}) - Azure CLI installed and authenticated (
az login) - Optional: Set
PROJECT_RESOURCE_IDenvironment variable
Workflow Overview
Complete Flow (14 Phases)
Fast Path (Defaults)
If user accepts all defaults (latest version, GlobalStandard SKU, recommended capacity, default RAI policy, standard upgrade policy), deployment completes in ~5 interactions.
Phase Summaries
⚠️ MUST READ: Before executing any phase, load references/customize-workflow.md [blocked] for the full scripts and implementation details. The summaries below describe what each phase does — the reference file contains the how (CLI commands, quota patterns, capacity formulas, cross-region fallback logic).
Error Handling
Common Issues and Resolutions
Troubleshooting Commands
Selection Guides & Advanced Topics
For SKU comparison tables, PTU sizing formulas, and advanced option details, load references/customize-guides.md [blocked].
SKU selection: GlobalStandard (production/HA) → Standard (dev/test) → ProvisionedManaged (high-volume/guaranteed throughput) → DataZoneStandard (data residency).
Capacity: TPM-based SKUs range from 1K (dev) to 100K+ (large production). PTU-based use formula: (Input TPM × 0.001) + (Output TPM × 0.002) + (Requests/min × 0.1).
Advanced options: Dynamic quota (GlobalStandard only), priority processing (PTU only, extra cost), spillover (overflow to backup deployment).
Related Skills
- preset - Quick deployment to best region with automatic configuration
- microsoft-foundry - Parent skill for all Microsoft Foundry operations
- quota — For quota viewing, increase requests, and troubleshooting quota errors, defer to this skill instead of duplicating guidance
- rbac - Manage permissions and access control
Notes
- Set
PROJECT_RESOURCE_IDenvironment variable to skip prompt - Not all SKUs available in all regions; capacity varies by subscription/region/model
- Custom RAI policies can be configured in Azure Portal
- Automatic version upgrades occur during maintenance windows
- Use Azure Monitor and Application Insights for production deployments


