IaC for Data Engineering with Terraform
Skill by ara.so — Data Skills collection.
This project provides Infrastructure-as-Code (IaC) templates and patterns for data engineers using Terraform to provision and manage AWS resources. It focuses on creating reproducible, version-controlled infrastructure for data platforms including S3 storage, EC2 compute instances, and IAM permissions.
What This Project Does
- Provides Terraform configurations for common data engineering infrastructure on AWS
- Demonstrates IaC best practices for S3 buckets, EC2 instances, and IAM roles
- Shows state management and lifecycle operations for data infrastructure
- Teaches reproducible infrastructure provisioning for data pipelines
Prerequisites
Before using this project, ensure you have:
- AWS Account with root or admin access
- Terraform CLI installed (installation guide)
- AWS CLI installed and configured (setup guide)
- AWS Credentials configured via
aws configure
AWS IAM Setup
Create an IAM user with appropriate permissions:
- Create IAM User: Navigate to AWS Console → IAM → Users → Create user
- Create Inline Policy: Attach a custom policy to the user
- Grant Permissions: For development/learning, grant full access to:
- Amazon S3
- Amazon EC2
- AWS IAM
⚠️ Security Note: Full service access is NOT recommended for production. Use least-privilege policies in production environments.
Project Structure
Key Terraform Commands
Initialize Terraform
Initialize the working directory and download provider plugins:
Validate Configuration
Check if the configuration is syntactically valid:
Format Code
Automatically format Terraform files to canonical style:
Plan Infrastructure Changes
Preview what Terraform will create/modify/destroy:
Apply Configuration
Create or update infrastructure:
Terraform will show a plan and ask for confirmation. Type yes to proceed.
Auto-approve (for automation)
Destroy Infrastructure
Remove all resources managed by Terraform:
Configuration
Basic Terraform Configuration Example
Before applying, modify terraform/main.tf to customize resource names:
Variables Configuration
Create terraform/variables.tf for reusable configurations:
Use variables in main.tf:
Create terraform/terraform.tfvars:
State Management
Inspect State
List all resources in the state:
View detailed state information:
Remote State (Production Pattern)
For production, store state remotely in S3:
Initialize with backend configuration:
Verification Commands
Verify S3 Bucket Creation
Verify EC2 Instance
Check Specific Resource
Common Patterns for Data Engineering
Pattern 1: Data Lake with Multiple Buckets
Pattern 2: EC2 with Data Processing Tools
Pattern 3: Outputs for Integration
Access outputs:
Troubleshooting
Issue: "Error acquiring the state lock"
Cause: Another Terraform process is running or a previous run didn't release the lock.
Solution:
Issue: "bucket name already exists"
Cause: S3 bucket names must be globally unique across all AWS accounts.
Solution: Change the bucket name in main.tf to something unique:
Issue: "insufficient IAM permissions"
Cause: The IAM user doesn't have required permissions.
Solution: Verify IAM policy includes necessary actions:
Issue: State file out of sync
Cause: Manual changes made outside Terraform.
Solution: Refresh the state:
Or import existing resources:
Workflow Example
Complete workflow for setting up data infrastructure:
Best Practices for Data Engineering IaC
- Use variables for environment-specific values
- Enable S3 versioning for data lineage and recovery
- Tag all resources for cost tracking and management
- Store state remotely in S3 with encryption and locking
- Use modules to organize reusable infrastructure components
- Never commit
.tfstatefiles or AWS credentials to version control - Implement lifecycle rules on S3 for cost optimization
- Use IAM roles instead of access keys for EC2 instances
- Plan before apply to review changes
- Destroy unused resources to avoid unnecessary costs


