Domino Distributed Computing Skill
Description
This skill helps users work with distributed computing frameworks in Domino - Spark, Ray, and Dask clusters for scaling compute-intensive workloads.
Activation
Activate this skill when users want to:
- Run Spark, Ray, or Dask clusters in Domino
- Scale data processing or ML training
- Configure distributed cluster settings
- Understand when to use each framework
Supported Frameworks
When to Use Each Framework
Spark
- Processing terabyte-scale data
- SQL analytics on big data
- ETL pipelines
- Structured data processing
Ray
- Distributed model training
- Hyperparameter optimization
- Reinforcement learning
- Generic Python parallelization
Dask
- Scaling pandas workflows
- Parallel NumPy operations
- Lazy evaluation needed
- Familiar pandas/NumPy API preferred
Launching On-Demand Clusters
Via Domino UI
- Start a workspace or job
- Check Attach compute cluster
- Select:
- Cluster Type: Spark, Ray, or Dask
- Worker Count: Number of workers
- Hardware Tier: Resources per worker
- Auto-scaling: Enable/disable
- Launch
Via Python SDK
Apache Spark
Connecting to Spark
Reading Data
Processing Data
Machine Learning with Spark MLlib
Writing Results
Ray
Connecting to Ray
Parallel Tasks
Distributed Training with Ray Train
Hyperparameter Tuning with Ray Tune
Dask
Connecting to Dask
Dask DataFrames (Parallel pandas)
Dask Arrays (Parallel NumPy)
Dask ML
GPU Clusters
Spark RAPIDS
Ray with GPUs
Autoscaling
Enable Autoscaling
Configure clusters to scale based on workload:
Monitor Scaling
View cluster status in Domino UI or via dashboard URLs.
Best Practices
1. Choose Right Framework
- SQL/ETL: Spark
- ML/Parallel Python: Ray
- Pandas at scale: Dask
2. Right-size Clusters
- Start small, scale up
- Monitor resource usage
- Use autoscaling when unsure
3. Data Locality
4. Persist Intermediate Results
Troubleshooting
Cluster Won't Start
- Check hardware tier availability
- Verify cluster environment builds
- Review cluster logs
Out of Memory
- Increase worker memory
- Add more workers
- Optimize code (reduce shuffles)
Slow Performance
- Check data locality
- Review partition sizes
- Monitor cluster dashboard


