SAP HANA ML Python Client (hana-ml)
Related Skills
- sap-dependency-security: Use for secure dependency pinning and upgrade workflows in Python/auxiliary tooling used alongside HANA ML stacks
When to Use This Skill
Use this skill when building machine learning workflows with the hana-ml Python client, using PAL/APL algorithms, querying HANA DataFrames, training or scoring models in-database, using AutoML, visualizing model output, or troubleshooting Python-to-HANA ML connections.
Common Issues
Package Version: 2.22.241011
Last Verified: 2025-11-27
Table of Contents
Installation & Setup
Requirements: Python 3.8+, SAP HANA 2.0 SPS03+ or SAP HANA Cloud
Quick Start
Connection & DataFrame
PAL Classification
APL AutoML
Model Persistence
Core Libraries
PAL (Predictive Analysis Library)
- 100+ algorithms executed in-database
- Categories: Classification, Regression, Clustering, Time Series, Preprocessing
- Key classes:
UnifiedClassification,UnifiedRegression,KMeans,ARIMA - See:
references/PAL_ALGORITHMS.mdfor complete list
APL (Automated Predictive Library)
- AutoML capabilities with automatic feature engineering
- Key classes:
AutoClassifier,AutoRegressor,GradientBoostingClassifier - See:
references/APL_ALGORITHMS.mdfor details
DataFrames
- Lazy evaluation - builds SQL until
collect()called - In-database processing for optimal performance
- See:
references/DATAFRAME_REFERENCE.mdfor complete API
Visualizers
- EDA plots, model explanations, metrics
- SHAP integration for model interpretability
- See:
references/VISUALIZERS.mdfor 14 visualization modules
Common Patterns
Train-Test Split
Feature Importance
Pipeline
Best Practices
- Use lazy evaluation - Operations build SQL without execution until
collect() - Leverage in-database processing - Keep data in HANA for performance
- Use Unified interfaces - Consistent APIs across algorithms
- Save models - Use
ModelStoragefor persistence - Explain predictions - Use SHAP explainers for interpretability
- Monitor AutoML - Use
PipelineProgressStatusMonitorfor long-running jobs
Bundled Resources
Reference Files
-
references/DATAFRAME_REFERENCE.md(479 lines)- ConnectionContext API, DataFrame operations, SQL generation
-
references/PAL_ALGORITHMS.md(869 lines)- Complete PAL algorithm reference (100+ algorithms)
- Classification, Regression, Clustering, Time Series, Preprocessing
-
references/APL_ALGORITHMS.md(534 lines)- AutoML capabilities, automated feature engineering
- AutoClassifier, AutoRegressor, GradientBoosting classes
-
references/VISUALIZERS.md(704 lines)- 14 visualization modules (EDA, SHAP, metrics, time series)
- Plot types, configuration, export options
-
references/SUPPORTING_MODULES.md(626 lines)- Model storage, spatial analytics, graph algorithms
- Text mining, statistics, error handling
Error Handling
Documentation
- Official Docs: https://help.sap.com/doc/1d0ebfe5e8dd44d09606814d83308d4b/2.0.07/en-US/hana_ml.html
- PyPI Package: https://pypi.org/project/hana-ml/


