Azure Storage File Datalake Py

作者 microsoft354361d83247MIT收錄於 2026年10月8日更新於 2026年10月8日

Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".

AI 產生的概覽

指導使用 Azure Data Lake Storage Gen2 的 Python SDK 進行檔案、目錄與存取控制操作。

功能
此技能說明如何使用 Azure Data Lake Storage Gen2 的 Python SDK,涵蓋以 DefaultAzureCredential 進行驗證、用戶端階層以及生命週期管理。它提供建立與刪除檔案系統、目錄和檔案,上傳與下載資料,列出路徑,讀取屬性以及設定 ACL 的程式碼範例。它也介紹非同步用戶端,並列出內容管理器、大型檔案使用 append/flush 等最佳實務。
適用情境
適用於撰寫針對 Azure Data Lake Storage Gen2 的 Python 程式碼,例如階層式檔案系統操作、大數據分析儲存,或檔案與目錄管理。適合涉及 DataLakeServiceClient、FileSystemClient 或 ADLS Gen2 階層式命名空間的工作。
執行需求
需要 Python 以及 azure-storage-file-datalake 和 azure-identity 套件、Azure 儲存體帳戶 URL,以及 DefaultAzureCredential 或受控識別等認證;需要連線至 Azure 的網路。此技能不含指令碼,只有說明與參考文件。

Azure Data Lake Storage Gen2 SDK for Python

Hierarchical file system for big data analytics workloads.

Installation

bash
pip install azure-storage-file-datalake azure-identity

Environment Variables

bash
AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net  # Required for all auth methodsAZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Prefer DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
    • Local dev: DefaultAzureCredential works as-is.
    • Production: set AZURE_TOKEN_CREDENTIALS=prod (or AZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.
  2. Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
    • Sync: with <Client>(...) as client:
    • Async: async with <Client>(...) as client: and async with DefaultAzureCredential() as credential: (from azure.identity.aio)

Snippets may abbreviate this setup, but production code should always follow both rules.

python
from azure.identity import DefaultAzureCredential, ManagedIdentityCredentialfrom azure.storage.filedatalake import DataLakeServiceClient
# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>credential = DefaultAzureCredential(require_envvar=True)# Or use a specific credential directly in production:# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes# credential = ManagedIdentityCredential()account_url = "https://<account>.dfs.core.windows.net"
with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client:    # Use service_client here (see following sections for operations)    ...

Client Hierarchy

ClientPurpose
DataLakeServiceClientAccount-level operations
FileSystemClientContainer (file system) operations
DataLakeDirectoryClientDirectory operations
DataLakeFileClientFile operations

File System Operations

python
# Create file system (container)file_system_client = service_client.create_file_system("myfilesystem")
# Get existingfile_system_client = service_client.get_file_system_client("myfilesystem")
# Deleteservice_client.delete_file_system("myfilesystem")
# List file systemsfor fs in service_client.list_file_systems():    print(fs.name)

Directory Operations

python
file_system_client = service_client.get_file_system_client("myfilesystem")
# Create directorydirectory_client = file_system_client.create_directory("mydir")
# Create nested directoriesdirectory_client = file_system_client.create_directory("path/to/nested/dir")
# Get directory clientdirectory_client = file_system_client.get_directory_client("mydir")
# Delete directorydirectory_client.delete_directory()
# Rename/move directorydirectory_client.rename_directory(new_name="myfilesystem/newname")

File Operations

Upload File

python
# Get file clientfile_client = file_system_client.get_file_client("path/to/file.txt")
# Upload from local filewith open("local-file.txt", "rb") as data:    file_client.upload_data(data, overwrite=True)
# Upload bytesfile_client.upload_data(b"Hello, Data Lake!", overwrite=True)
# Append data (for large files)file_client.append_data(data=b"chunk1", offset=0, length=6)file_client.append_data(data=b"chunk2", offset=6, length=6)file_client.flush_data(12)  # Commit the data

Download File

python
file_client = file_system_client.get_file_client("path/to/file.txt")
# Download all contentdownload = file_client.download_file()content = download.readall()
# Download to filewith open("downloaded.txt", "wb") as f:    download = file_client.download_file()    download.readinto(f)
# Download rangedownload = file_client.download_file(offset=0, length=100)

Delete File

python
file_client.delete_file()

List Contents

python
# List paths (files and directories)for path in file_system_client.get_paths():    print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}")
# List paths in directoryfor path in file_system_client.get_paths(path="mydir"):    print(path.name)
# Recursive listingfor path in file_system_client.get_paths(path="mydir", recursive=True):    print(path.name)

File/Directory Properties

python
# Get propertiesproperties = file_client.get_file_properties()print(f"Size: {properties.size}")print(f"Last modified: {properties.last_modified}")
# Set metadatafile_client.set_metadata(metadata={"processed": "true"})

Access Control (ACL)

python
# Get ACLacl = directory_client.get_access_control()print(f"Owner: {acl['owner']}")print(f"Permissions: {acl['permissions']}")
# Set ACLdirectory_client.set_access_control(    owner="user-id",    permissions="rwxr-x---")
# Update ACL entriesfrom azure.storage.filedatalake import AccessControlChangeResultdirectory_client.update_access_control_recursive(    acl="user:user-id:rwx")

Async Client

python
from azure.storage.filedatalake.aio import DataLakeServiceClientfrom azure.identity.aio import DefaultAzureCredential
async def datalake_operations():    async with DefaultAzureCredential() as credential:        async with DataLakeServiceClient(            account_url="https://<account>.dfs.core.windows.net",            credential=credential        ) as service_client:            file_system_client = service_client.get_file_system_client("myfilesystem")            file_client = file_system_client.get_file_client("test.txt")                        await file_client.upload_data(b"async content", overwrite=True)                        download = await file_client.download_file()            content = await download.readall()
import asyncioasyncio.run(datalake_operations())

Best Practices

  1. Pick sync OR async and stay consistent. Do not mix azure.storage.filedatalake sync clients with azure.storage.filedatalake.aio async clients in the same call path. Choose one mode per module.
  2. Always use context managers for clients and async credentials. Wrap every client in with DataLakeServiceClient(...) as client: (sync) or async with DataLakeServiceClient(...) as client: (async). For async DefaultAzureCredential from azure.identity.aio, also use async with credential: so tokens and transports are cleaned up.
  3. Use DefaultAzureCredential for portable auth across local dev and Azure (avoid connection strings / API keys when possible).
  4. Use hierarchical namespace for file system semantics
  5. Use append_data + flush_data for large file uploads
  6. Set ACLs at directory level and inherit to children
  7. Use async client for high-throughput scenarios
  8. Use get_paths with recursive=True for full directory listing
  9. Set metadata for custom file attributes
  10. Consider Blob API for simple object storage use cases

Reference Files

FileContents
references/capabilities.md [blocked]Additional non-hero capabilities, operation-group coverage, and production checklists.
references/non-hero-scenarios.md [blocked]Dedicated non-hero examples for secondary/advanced scenarios.

來源與署名

來源:microsoft/skills位於.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py提交354361d

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架