Domino Datasets Skill
Description
This skill helps users work with Domino Datasets - high-performance, versioned filesystem storage for data science projects.
Activation
Activate this skill when users want to:
- Create or manage Domino Datasets
- Work with dataset snapshots and versioning
- Share data between projects
- Access large datasets efficiently
- Understand dataset paths and mounting
What is a Domino Dataset?
A Domino Dataset is:
- High-performance storage: Network filesystem optimized for data science
- Versioned: Create snapshots for reproducibility
- Shareable: Access across projects
- Scalable: No file size or count limits
- Persistent: Data persists across executions
Creating a Dataset
Via Domino UI
- Navigate to your project
- Go to Data > Domino Datasets
- Click Create New Dataset
- Enter:
- Name: Dataset name (e.g.,
training-data) - Description: What the dataset contains
- Name: Dataset name (e.g.,
- Click Create
Via Python SDK
Dataset Paths
Dataset paths differ based on your project type. Domino has two project types with different mount structures.
DFS (Domino File System) Projects
DFS projects use /domino as the root:
Git-Based Projects
Git-based projects use /mnt as the root:
How to Identify Your Project Type
Check which paths exist in your execution:
Permissions
Both project types follow the same permission model:
- Owners/Editors: Read-write access to datasets
- Readers: Read-only access
Example: Reading Data
Uploading Data
Via Domino UI
- Go to dataset page
- Click Upload
- Select files (up to 50GB or 50,000 files via UI)
- Click Upload
Via Domino CLI (Large Uploads)
Via Code in Workspace
Snapshots
What is a Snapshot?
A snapshot is a read-only, immutable version of your dataset at a point in time. Use snapshots for:
- Reproducibility
- Versioning training data
- Rolling back to previous states
Create a Snapshot
Or via UI:
- Go to dataset page
- Click Create Snapshot
- Add optional tag (e.g.,
v1.0,production)
Access Snapshots
Snapshot Limits
- Default limit: 20 snapshots per dataset
- Configurable by admins
- Oldest snapshots auto-deleted when limit reached
Tags
What are Tags?
Tags provide friendly names for snapshots:
production: Current production datav1.0,v2.0: Version numbers2024-01-15: Date-based tags
Move Tags
Tags can be moved to different snapshots:
Sharing Datasets
Within Organization
- Go to dataset settings
- Set visibility to Organization
- Other projects can mount the dataset
Cross-Project Access
Best Practices
1. Use Appropriate Storage
2. Organize Data
3. Use Efficient Formats
4. Document Data
Include README and schema:
5. Snapshot Before Changes
Reading Large Datasets
Chunked Reading
Lazy Loading with Dask
Memory Mapping
Troubleshooting
Dataset Not Found
- Verify dataset name is correct
- Check dataset is mounted to project
- Confirm you have access permissions
Permission Denied
- Check project role (need Contributor+)
- Verify dataset sharing settings
- Contact dataset owner
Slow Performance
- Use efficient file formats (Parquet > CSV)
- Read only needed columns
- Use chunked/lazy loading for large files
Snapshot Failed
- Check disk quota
- Verify no files are open/locked
- Check snapshot limit not reached

