Azure Storage File Datalake Py

by microsoft354361d83247MITListed Oct 8, 2026Updated Oct 8, 2026

Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".

FeaturedInstructions onlySoftware DevelopmentDevOps & Cloud
AI-generated overview

Guides Python use of the Azure Data Lake Storage Gen2 SDK for file, directory, and access-control operations.

What it does
This skill documents how to use the Azure Data Lake Storage Gen2 SDK for Python, covering authentication with DefaultAzureCredential, client hierarchy, and lifecycle management. It shows code for creating and deleting file systems, directories, and files, uploading and downloading data, listing paths, reading properties, and setting ACLs. It also covers the async client and lists best practices such as context managers and append/flush for large uploads.
When to use it
Use it when writing Python code against Azure Data Lake Storage Gen2, such as hierarchical file system operations, big data analytics storage, or file and directory management. It suits tasks involving DataLakeServiceClient, FileSystemClient, or ADLS Gen2 hierarchical namespaces.
Requirements
Requires Python with the azure-storage-file-datalake and azure-identity packages, an Azure Storage account URL, and credentials such as DefaultAzureCredential or managed identity; network access to Azure is needed. It ships no scripts, only instructions and reference documents.

Azure Data Lake Storage Gen2 SDK for Python

Hierarchical file system for big data analytics workloads.

Installation

bash
pip install azure-storage-file-datalake azure-identity

Environment Variables

bash
AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net  # Required for all auth methodsAZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Prefer DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
    • Local dev: DefaultAzureCredential works as-is.
    • Production: set AZURE_TOKEN_CREDENTIALS=prod (or AZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.
  2. Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
    • Sync: with <Client>(...) as client:
    • Async: async with <Client>(...) as client: and async with DefaultAzureCredential() as credential: (from azure.identity.aio)

Snippets may abbreviate this setup, but production code should always follow both rules.

python
from azure.identity import DefaultAzureCredential, ManagedIdentityCredentialfrom azure.storage.filedatalake import DataLakeServiceClient
# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>credential = DefaultAzureCredential(require_envvar=True)# Or use a specific credential directly in production:# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes# credential = ManagedIdentityCredential()account_url = "https://<account>.dfs.core.windows.net"
with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client:    # Use service_client here (see following sections for operations)    ...

Client Hierarchy

ClientPurpose
DataLakeServiceClientAccount-level operations
FileSystemClientContainer (file system) operations
DataLakeDirectoryClientDirectory operations
DataLakeFileClientFile operations

File System Operations

python
# Create file system (container)file_system_client = service_client.create_file_system("myfilesystem")
# Get existingfile_system_client = service_client.get_file_system_client("myfilesystem")
# Deleteservice_client.delete_file_system("myfilesystem")
# List file systemsfor fs in service_client.list_file_systems():    print(fs.name)

Directory Operations

python
file_system_client = service_client.get_file_system_client("myfilesystem")
# Create directorydirectory_client = file_system_client.create_directory("mydir")
# Create nested directoriesdirectory_client = file_system_client.create_directory("path/to/nested/dir")
# Get directory clientdirectory_client = file_system_client.get_directory_client("mydir")
# Delete directorydirectory_client.delete_directory()
# Rename/move directorydirectory_client.rename_directory(new_name="myfilesystem/newname")

File Operations

Upload File

python
# Get file clientfile_client = file_system_client.get_file_client("path/to/file.txt")
# Upload from local filewith open("local-file.txt", "rb") as data:    file_client.upload_data(data, overwrite=True)
# Upload bytesfile_client.upload_data(b"Hello, Data Lake!", overwrite=True)
# Append data (for large files)file_client.append_data(data=b"chunk1", offset=0, length=6)file_client.append_data(data=b"chunk2", offset=6, length=6)file_client.flush_data(12)  # Commit the data

Download File

python
file_client = file_system_client.get_file_client("path/to/file.txt")
# Download all contentdownload = file_client.download_file()content = download.readall()
# Download to filewith open("downloaded.txt", "wb") as f:    download = file_client.download_file()    download.readinto(f)
# Download rangedownload = file_client.download_file(offset=0, length=100)

Delete File

python
file_client.delete_file()

List Contents

python
# List paths (files and directories)for path in file_system_client.get_paths():    print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}")
# List paths in directoryfor path in file_system_client.get_paths(path="mydir"):    print(path.name)
# Recursive listingfor path in file_system_client.get_paths(path="mydir", recursive=True):    print(path.name)

File/Directory Properties

python
# Get propertiesproperties = file_client.get_file_properties()print(f"Size: {properties.size}")print(f"Last modified: {properties.last_modified}")
# Set metadatafile_client.set_metadata(metadata={"processed": "true"})

Access Control (ACL)

python
# Get ACLacl = directory_client.get_access_control()print(f"Owner: {acl['owner']}")print(f"Permissions: {acl['permissions']}")
# Set ACLdirectory_client.set_access_control(    owner="user-id",    permissions="rwxr-x---")
# Update ACL entriesfrom azure.storage.filedatalake import AccessControlChangeResultdirectory_client.update_access_control_recursive(    acl="user:user-id:rwx")

Async Client

python
from azure.storage.filedatalake.aio import DataLakeServiceClientfrom azure.identity.aio import DefaultAzureCredential
async def datalake_operations():    async with DefaultAzureCredential() as credential:        async with DataLakeServiceClient(            account_url="https://<account>.dfs.core.windows.net",            credential=credential        ) as service_client:            file_system_client = service_client.get_file_system_client("myfilesystem")            file_client = file_system_client.get_file_client("test.txt")                        await file_client.upload_data(b"async content", overwrite=True)                        download = await file_client.download_file()            content = await download.readall()
import asyncioasyncio.run(datalake_operations())

Best Practices

  1. Pick sync OR async and stay consistent. Do not mix azure.storage.filedatalake sync clients with azure.storage.filedatalake.aio async clients in the same call path. Choose one mode per module.
  2. Always use context managers for clients and async credentials. Wrap every client in with DataLakeServiceClient(...) as client: (sync) or async with DataLakeServiceClient(...) as client: (async). For async DefaultAzureCredential from azure.identity.aio, also use async with credential: so tokens and transports are cleaned up.
  3. Use DefaultAzureCredential for portable auth across local dev and Azure (avoid connection strings / API keys when possible).
  4. Use hierarchical namespace for file system semantics
  5. Use append_data + flush_data for large file uploads
  6. Set ACLs at directory level and inherit to children
  7. Use async client for high-throughput scenarios
  8. Use get_paths with recursive=True for full directory listing
  9. Set metadata for custom file attributes
  10. Consider Blob API for simple object storage use cases

Reference Files

FileContents
references/capabilities.md [blocked]Additional non-hero capabilities, operation-group coverage, and production checklists.
references/non-hero-scenarios.md [blocked]Dedicated non-hero examples for secondary/advanced scenarios.

Source and attribution

Source:microsoft/skillsin.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-pyat commit354361d

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal