Sap Hana Ml

secondsky/sap-skills/plugins/sap-hana-ml/skills/sap-hana-ml

作者 secondsky652a861d3ed422b71f01c8805f9ec03b14017cd8GPL-3.0收录于 2026年10月9日更新于 2026年10月9日

SAP HANA Machine Learning Python Client (hana-ml) development skill. Use when: Building ML solutions with SAP HANA's in-database machine learning using Python hana-ml library for PAL/APL algorithms, DataFrame operations, AutoML, model persistence, and visualization. Keywords: hana-ml, SAP HANA, machine learning, PAL, APL, predictive analytics, HANA DataFrame, ConnectionContext, classification, regression, clustering, time series, ARIMA, gradient boosting, AutoML, SHAP, model storage

AI 生成的概览

指导使用 SAP HANA 的 hana-ml Python 客户端进行库内机器学习开发,涵盖 PAL/APL 算法、DataFrame、AutoML 与模型存储。

功能
该技能提供使用 hana-ml Python 客户端在 SAP HANA 上构建机器学习工作流的说明与参考资料。内容涵盖通过 ConnectionContext 建立连接、以惰性求值方式查询 HANA DataFrame、训练与评分 PAL 和 APL 模型、AutoML、管道、使用 ModelStorage 持久化模型,以及包括 SHAP 解释在内的可视化。它产出的是代码模式与指导,而非文件,且不附带脚本。
适用场景
适用于使用 hana-ml Python 库构建或排查在 SAP HANA 内部运行的机器学习解决方案。适合涉及 PAL 或 APL 算法、HANA DataFrame、AutoML、模型存储或 Python 与 HANA 连接问题的任务。
运行要求
需要 hana-ml Python 包(文中提及版本 2.22.241011)、Python 3.8+,以及可访问已安装并授权所需 AFL/PAL/APL 库的 SAP HANA 2.0 SPS03+ 或 SAP HANA Cloud。连接需要 HANA 主机、端口、用户凭据、TLS/加密和网络访问。不附带脚本,仅包含说明与参考文档。

SAP HANA ML Python Client (hana-ml)

Related Skills

  • sap-dependency-security: Use for secure dependency pinning and upgrade workflows in Python/auxiliary tooling used alongside HANA ML stacks

When to Use This Skill

Use this skill when building machine learning workflows with the hana-ml Python client, using PAL/APL algorithms, querying HANA DataFrames, training or scoring models in-database, using AutoML, visualizing model output, or troubleshooting Python-to-HANA ML connections.

Common Issues

IssueFirst check
Connection failsVerify HANA host, port, TLS/encryption, user privileges, and network allowlists.
PAL/APL algorithm missingConfirm the HANA system has the required AFL/PAL/APL libraries installed and licensed.
DataFrame collection is slowPush filtering/projection into HANA and avoid collecting large frames into Python.

Package Version: 2.22.241011
Last Verified: 2025-11-27

Table of Contents


Installation & Setup

bash
pip install hana-ml

Requirements: Python 3.8+, SAP HANA 2.0 SPS03+ or SAP HANA Cloud


Quick Start

Connection & DataFrame

python
from hana_ml import ConnectionContext
# Connectconn = ConnectionContext(    address='<hostname>',    port=443,    user='<username>',    password='<password>',    encrypt=True)
# Create DataFramedf = conn.table('MY_TABLE', schema='MY_SCHEMA')print(f"Shape: {df.shape}")df.head(10).collect()

PAL Classification

python
from hana_ml.algorithms.pal.unified_classification import UnifiedClassification
# Train modelclf = UnifiedClassification(func='RandomDecisionTree')clf.fit(train_df, features=['F1', 'F2', 'F3'], label='TARGET')
# Predict & evaluatepredictions = clf.predict(test_df, features=['F1', 'F2', 'F3'])score = clf.score(test_df, features=['F1', 'F2', 'F3'], label='TARGET')

APL AutoML

python
from hana_ml.algorithms.apl.classification import AutoClassifier
# Automated classificationauto_clf = AutoClassifier()auto_clf.fit(train_df, label='TARGET')predictions = auto_clf.predict(test_df)

Model Persistence

python
from hana_ml.model_storage import ModelStorage
ms = ModelStorage(conn)clf.name = 'MY_CLASSIFIER'ms.save_model(model=clf, if_exists='replace')

Core Libraries

PAL (Predictive Analysis Library)

  • 100+ algorithms executed in-database
  • Categories: Classification, Regression, Clustering, Time Series, Preprocessing
  • Key classes: UnifiedClassification, UnifiedRegression, KMeans, ARIMA
  • See: references/PAL_ALGORITHMS.md for complete list

APL (Automated Predictive Library)

  • AutoML capabilities with automatic feature engineering
  • Key classes: AutoClassifier, AutoRegressor, GradientBoostingClassifier
  • See: references/APL_ALGORITHMS.md for details

DataFrames

  • Lazy evaluation - builds SQL until collect() called
  • In-database processing for optimal performance
  • See: references/DATAFRAME_REFERENCE.md for complete API

Visualizers

  • EDA plots, model explanations, metrics
  • SHAP integration for model interpretability
  • See: references/VISUALIZERS.md for 14 visualization modules

Common Patterns

Train-Test Split

python
from hana_ml.algorithms.pal.partition import train_test_val_split
train, test, val = train_test_val_split(    data=df,    training_percentage=0.7,    testing_percentage=0.2,    validation_percentage=0.1)

Feature Importance

python
# APL modelsimportance = auto_clf.get_feature_importances()
# PAL modelsfrom hana_ml.algorithms.pal.preprocessing import FeatureSelectionfs = FeatureSelection()fs.fit(train_df, features=features, label='TARGET')

Pipeline

python
from hana_ml.algorithms.pal.pipeline import Pipelinefrom hana_ml.algorithms.pal.preprocessing import Imputer, FeatureNormalizer
pipeline = Pipeline([    ('imputer', Imputer(strategy='mean')),    ('normalizer', FeatureNormalizer()),    ('classifier', UnifiedClassification(func='RandomDecisionTree'))])

Best Practices

  1. Use lazy evaluation - Operations build SQL without execution until collect()
  2. Leverage in-database processing - Keep data in HANA for performance
  3. Use Unified interfaces - Consistent APIs across algorithms
  4. Save models - Use ModelStorage for persistence
  5. Explain predictions - Use SHAP explainers for interpretability
  6. Monitor AutoML - Use PipelineProgressStatusMonitor for long-running jobs

Bundled Resources

Reference Files

  • references/DATAFRAME_REFERENCE.md (479 lines)

    • ConnectionContext API, DataFrame operations, SQL generation
  • references/PAL_ALGORITHMS.md (869 lines)

    • Complete PAL algorithm reference (100+ algorithms)
    • Classification, Regression, Clustering, Time Series, Preprocessing
  • references/APL_ALGORITHMS.md (534 lines)

    • AutoML capabilities, automated feature engineering
    • AutoClassifier, AutoRegressor, GradientBoosting classes
  • references/VISUALIZERS.md (704 lines)

    • 14 visualization modules (EDA, SHAP, metrics, time series)
    • Plot types, configuration, export options
  • references/SUPPORTING_MODULES.md (626 lines)

    • Model storage, spatial analytics, graph algorithms
    • Text mining, statistics, error handling

Error Handling

python
from hana_ml.ml_exceptions import Error
try:    clf.fit(train_df, features=features, label='TARGET')except Error as e:    print(f"HANA ML Error: {e}")

Documentation

来源与署名

来源:secondsky/sap-skills位于plugins/sap-hana-ml/skills/sap-hana-ml提交652a861

许可证: GPL-3.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架