Azure Storage File Datalake Py

作者 microsoft354361d83247MIT收录于 2026年10月8日更新于 2026年10月8日

Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".

AI 生成的概览

指导使用 Azure Data Lake Storage Gen2 的 Python SDK 进行文件、目录和访问控制操作。

功能
该技能说明如何使用 Azure Data Lake Storage Gen2 的 Python SDK,涵盖使用 DefaultAzureCredential 进行身份验证、客户端层级以及生命周期管理。它给出创建和删除文件系统、目录与文件,上传和下载数据,列出路径,读取属性以及设置 ACL 的代码示例。它还介绍异步客户端,并列出上下文管理器、大文件使用 append/flush 等最佳实践。
适用场景
适用于编写针对 Azure Data Lake Storage Gen2 的 Python 代码,例如分层文件系统操作、大数据分析存储或文件与目录管理。适合涉及 DataLakeServiceClient、FileSystemClient 或 ADLS Gen2 分层命名空间的任务。
运行要求
需要 Python 以及 azure-storage-file-datalake 和 azure-identity 包、Azure 存储账户 URL,以及 DefaultAzureCredential 或托管标识等凭据;需要访问 Azure 的网络。该技能不包含脚本,只有说明和参考文档。

Azure Data Lake Storage Gen2 SDK for Python

Hierarchical file system for big data analytics workloads.

Installation

bash
pip install azure-storage-file-datalake azure-identity

Environment Variables

bash
AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net  # Required for all auth methodsAZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Prefer DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
    • Local dev: DefaultAzureCredential works as-is.
    • Production: set AZURE_TOKEN_CREDENTIALS=prod (or AZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.
  2. Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
    • Sync: with <Client>(...) as client:
    • Async: async with <Client>(...) as client: and async with DefaultAzureCredential() as credential: (from azure.identity.aio)

Snippets may abbreviate this setup, but production code should always follow both rules.

python
from azure.identity import DefaultAzureCredential, ManagedIdentityCredentialfrom azure.storage.filedatalake import DataLakeServiceClient
# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>credential = DefaultAzureCredential(require_envvar=True)# Or use a specific credential directly in production:# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes# credential = ManagedIdentityCredential()account_url = "https://<account>.dfs.core.windows.net"
with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client:    # Use service_client here (see following sections for operations)    ...

Client Hierarchy

ClientPurpose
DataLakeServiceClientAccount-level operations
FileSystemClientContainer (file system) operations
DataLakeDirectoryClientDirectory operations
DataLakeFileClientFile operations

File System Operations

python
# Create file system (container)file_system_client = service_client.create_file_system("myfilesystem")
# Get existingfile_system_client = service_client.get_file_system_client("myfilesystem")
# Deleteservice_client.delete_file_system("myfilesystem")
# List file systemsfor fs in service_client.list_file_systems():    print(fs.name)

Directory Operations

python
file_system_client = service_client.get_file_system_client("myfilesystem")
# Create directorydirectory_client = file_system_client.create_directory("mydir")
# Create nested directoriesdirectory_client = file_system_client.create_directory("path/to/nested/dir")
# Get directory clientdirectory_client = file_system_client.get_directory_client("mydir")
# Delete directorydirectory_client.delete_directory()
# Rename/move directorydirectory_client.rename_directory(new_name="myfilesystem/newname")

File Operations

Upload File

python
# Get file clientfile_client = file_system_client.get_file_client("path/to/file.txt")
# Upload from local filewith open("local-file.txt", "rb") as data:    file_client.upload_data(data, overwrite=True)
# Upload bytesfile_client.upload_data(b"Hello, Data Lake!", overwrite=True)
# Append data (for large files)file_client.append_data(data=b"chunk1", offset=0, length=6)file_client.append_data(data=b"chunk2", offset=6, length=6)file_client.flush_data(12)  # Commit the data

Download File

python
file_client = file_system_client.get_file_client("path/to/file.txt")
# Download all contentdownload = file_client.download_file()content = download.readall()
# Download to filewith open("downloaded.txt", "wb") as f:    download = file_client.download_file()    download.readinto(f)
# Download rangedownload = file_client.download_file(offset=0, length=100)

Delete File

python
file_client.delete_file()

List Contents

python
# List paths (files and directories)for path in file_system_client.get_paths():    print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}")
# List paths in directoryfor path in file_system_client.get_paths(path="mydir"):    print(path.name)
# Recursive listingfor path in file_system_client.get_paths(path="mydir", recursive=True):    print(path.name)

File/Directory Properties

python
# Get propertiesproperties = file_client.get_file_properties()print(f"Size: {properties.size}")print(f"Last modified: {properties.last_modified}")
# Set metadatafile_client.set_metadata(metadata={"processed": "true"})

Access Control (ACL)

python
# Get ACLacl = directory_client.get_access_control()print(f"Owner: {acl['owner']}")print(f"Permissions: {acl['permissions']}")
# Set ACLdirectory_client.set_access_control(    owner="user-id",    permissions="rwxr-x---")
# Update ACL entriesfrom azure.storage.filedatalake import AccessControlChangeResultdirectory_client.update_access_control_recursive(    acl="user:user-id:rwx")

Async Client

python
from azure.storage.filedatalake.aio import DataLakeServiceClientfrom azure.identity.aio import DefaultAzureCredential
async def datalake_operations():    async with DefaultAzureCredential() as credential:        async with DataLakeServiceClient(            account_url="https://<account>.dfs.core.windows.net",            credential=credential        ) as service_client:            file_system_client = service_client.get_file_system_client("myfilesystem")            file_client = file_system_client.get_file_client("test.txt")                        await file_client.upload_data(b"async content", overwrite=True)                        download = await file_client.download_file()            content = await download.readall()
import asyncioasyncio.run(datalake_operations())

Best Practices

  1. Pick sync OR async and stay consistent. Do not mix azure.storage.filedatalake sync clients with azure.storage.filedatalake.aio async clients in the same call path. Choose one mode per module.
  2. Always use context managers for clients and async credentials. Wrap every client in with DataLakeServiceClient(...) as client: (sync) or async with DataLakeServiceClient(...) as client: (async). For async DefaultAzureCredential from azure.identity.aio, also use async with credential: so tokens and transports are cleaned up.
  3. Use DefaultAzureCredential for portable auth across local dev and Azure (avoid connection strings / API keys when possible).
  4. Use hierarchical namespace for file system semantics
  5. Use append_data + flush_data for large file uploads
  6. Set ACLs at directory level and inherit to children
  7. Use async client for high-throughput scenarios
  8. Use get_paths with recursive=True for full directory listing
  9. Set metadata for custom file attributes
  10. Consider Blob API for simple object storage use cases

Reference Files

FileContents
references/capabilities.md [blocked]Additional non-hero capabilities, operation-group coverage, and production checklists.
references/non-hero-scenarios.md [blocked]Dedicated non-hero examples for secondary/advanced scenarios.

来源与署名

来源:microsoft/skills位于.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py提交354361d

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架