Official agent skill

Azure Storage File Datalake Py

by microsoft in microsoft/skills

Azure Data Lake Storage Gen2 SDK for Python. An agent skill from microsoft/skills.

OfficialMITAuto-check passedDevOps & Cloud

Install Azure Storage File Datalake Py

skills CLI
$ npx skills add microsoft/skills --skill azure-storage-file-datalake-py -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install microsoft/skills azure-storage-file-datalake-py --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/microsoft/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py .claude/skills/azure-storage-file-datalake-py && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
azure-storage-file-datalake-py
GitHub stars
3.1k
Token cost
~2k tokens
SKILL.md length
341 words
Files
3 (incl. references)
Skills in repo
150
Repo updated
First seen
Licence
MIT

At a glance

Azure Data Lake Storage Gen2 SDK for Python. An agent skill from microsoft/skills.

  • Works in 10 steps: Pick sync OR async and stay consistent.… → Always use context managers for clients… → Use DefaultAzureCredential for portable… → …
  • Hierarchical file systems
  • SKILL.md covers Installation, Environment Variables, Authentication & Lifecycle and Client Hierarchy, plus 9 more sections
  • Calls pip; reaches learn.microsoft.com; needs AZURE_TOKEN_CREDENTIALS

What it does

Azure Storage File Datalake Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/capabilities.md` and `references/non-hero-scenarios.md`).

It sits in DevOps & Cloud, covering Data analysis. It works with Microsoft Azure, Python and Visual Studio Code. The repository describes itself as: Skills, MCP servers, Custom Agents, Agents.md for SDKs to ground Coding Agents. The licence is MIT.

When your agent uses it

  • Hierarchical file systems
  • Big data analytics
  • File/directory operations

Example prompts

  • “data lake”
  • “DataLakeServiceClient”
  • “FileSystemClient”
  • “/azure-storage-file-datalake-py”

Requirements

  • Python 3

Workflow steps

10 steps, taken from the first numbered list in SKILL.md.

  1. Pick sync OR async and stay consistent. Do not mix azure.storage.filedatalake sync clients with azure.storage.filedatalake.aio async…
  2. Always use context managers for clients and async credentials. Wrap every client in with DataLakeServiceClient(...) as client: (sync) or…
  3. Use DefaultAzureCredential for portable auth across local dev and Azure (avoid connection strings / API keys when possible).
  4. Use hierarchical namespace for file system semantics
  5. Use append_data + flush_data for large file uploads
  6. Set ACLs at directory level and inherit to children
  7. Use async client for high-throughput scenarios
  8. Use get_paths with recursive=True for full directory listing
  9. Set metadata for custom file attributes
  10. Consider Blob API for simple object storage use cases

What it can do on your machine

Read from SKILL.md and the folder at commit 3898ec8. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • learn.microsoft.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • AZURE_TOKEN_CREDENTIALS

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Azure Storage File Datalake Py loads about 2k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 67 tokens; SKILL.md has 341 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from microsoft/skills at commit 3898ec8, republished under its MIT licence (© microsoft). 341 words, ~2,043 tokens.

Download SKILL.mdSave it as .claude/skills/azure-storage-file-datalake-py/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
azure-storage-file-datalake-py
description
Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations. Triggers: "data lake", "DataLakeServiceClient", "FileSystemClient", "ADLS Gen2", "hierarchical namespace".
license
MIT
metadata.author
Microsoft
metadata.version
1.0.0
metadata.package
azure-storage-file-datalake

Azure Data Lake Storage Gen2 SDK for Python

Hierarchical file system for big data analytics workloads.

Installation

bash
pip install azure-storage-file-datalake azure-identity

Environment Variables

bash
AZURE_STORAGE_ACCOUNT_URL=https://<account>.dfs.core.windows.net  # Required for all auth methods
AZURE_TOKEN_CREDENTIALS=prod # Required only if DefaultAzureCredential is used in production

Authentication & Lifecycle

🔑 Two rules apply to every code sample below:

  1. Prefer DefaultAzureCredential. It works locally (Azure CLI / VS Code / Developer CLI) and in Azure (managed identity, workload identity) with no code change. Avoid connection strings, account/API keys — they bypass Entra audit and rotation.
    • Local dev: DefaultAzureCredential works as-is.
    • Production: set AZURE_TOKEN_CREDENTIALS=prod (or AZURE_TOKEN_CREDENTIALS=<specific_credential>) to constrain the credential chain to production-safe credentials.
  2. Wrap every client in a context manager so HTTP transports, sockets, and token caches are released deterministically:
    • Sync: with <Client>(...) as client:
    • Async: async with <Client>(...) as client: and async with DefaultAzureCredential() as credential: (from azure.identity.aio)

Snippets may abbreviate this setup, but production code should always follow both rules.

python
from azure.identity import DefaultAzureCredential, ManagedIdentityCredential
from azure.storage.filedatalake import DataLakeServiceClient

# Local dev: DefaultAzureCredential. Production: set AZURE_TOKEN_CREDENTIALS=prod or AZURE_TOKEN_CREDENTIALS=<specific_credential>
credential = DefaultAzureCredential(require_envvar=True)
# Or use a specific credential directly in production:
# See https://learn.microsoft.com/python/api/overview/azure/identity-readme?view=azure-python#credential-classes
# credential = ManagedIdentityCredential()
account_url = "https://<account>.dfs.core.windows.net"

with DataLakeServiceClient(account_url=account_url, credential=credential) as service_client:
    # Use service_client here (see following sections for operations)
    ...

Client Hierarchy

ClientPurpose
DataLakeServiceClientAccount-level operations
FileSystemClientContainer (file system) operations
DataLakeDirectoryClientDirectory operations
DataLakeFileClientFile operations

File System Operations

python
# Create file system (container)
file_system_client = service_client.create_file_system("myfilesystem")

# Get existing
file_system_client = service_client.get_file_system_client("myfilesystem")

# Delete
service_client.delete_file_system("myfilesystem")

# List file systems
for fs in service_client.list_file_systems():
    print(fs.name)

Directory Operations

python
file_system_client = service_client.get_file_system_client("myfilesystem")

# Create directory
directory_client = file_system_client.create_directory("mydir")

# Create nested directories
directory_client = file_system_client.create_directory("path/to/nested/dir")

# Get directory client
directory_client = file_system_client.get_directory_client("mydir")

# Delete directory
directory_client.delete_directory()

# Rename/move directory
directory_client.rename_directory(new_name="myfilesystem/newname")

File Operations

Upload File
python
# Get file client
file_client = file_system_client.get_file_client("path/to/file.txt")

# Upload from local file
with open("local-file.txt", "rb") as data:
    file_client.upload_data(data, overwrite=True)

# Upload bytes
file_client.upload_data(b"Hello, Data Lake!", overwrite=True)

# Append data (for large files)
file_client.append_data(data=b"chunk1", offset=0, length=6)
file_client.append_data(data=b"chunk2", offset=6, length=6)
file_client.flush_data(12)  # Commit the data
Download File
python
file_client = file_system_client.get_file_client("path/to/file.txt")

# Download all content
download = file_client.download_file()
content = download.readall()

# Download to file
with open("downloaded.txt", "wb") as f:
    download = file_client.download_file()
    download.readinto(f)

# Download range
download = file_client.download_file(offset=0, length=100)
Delete File
python
file_client.delete_file()

List Contents

python
# List paths (files and directories)
for path in file_system_client.get_paths():
    print(f"{'DIR' if path.is_directory else 'FILE'}: {path.name}")

# List paths in directory
for path in file_system_client.get_paths(path="mydir"):
    print(path.name)

# Recursive listing
for path in file_system_client.get_paths(path="mydir", recursive=True):
    print(path.name)

File/Directory Properties

python
# Get properties
properties = file_client.get_file_properties()
print(f"Size: {properties.size}")
print(f"Last modified: {properties.last_modified}")

# Set metadata
file_client.set_metadata(metadata={"processed": "true"})

Access Control (ACL)

python
# Get ACL
acl = directory_client.get_access_control()
print(f"Owner: {acl['owner']}")
print(f"Permissions: {acl['permissions']}")

# Set ACL
directory_client.set_access_control(
    owner="user-id",
    permissions="rwxr-x---"
)

# Update ACL entries
from azure.storage.filedatalake import AccessControlChangeResult
directory_client.update_access_control_recursive(
    acl="user:user-id:rwx"
)

Async Client

python
from azure.storage.filedatalake.aio import DataLakeServiceClient
from azure.identity.aio import DefaultAzureCredential

async def datalake_operations():
    async with DefaultAzureCredential() as credential:
        async with DataLakeServiceClient(
            account_url="https://<account>.dfs.core.windows.net",
            credential=credential
        ) as service_client:
            file_system_client = service_client.get_file_system_client("myfilesystem")
            file_client = file_system_client.get_file_client("test.txt")
            
            await file_client.upload_data(b"async content", overwrite=True)
            
            download = await file_client.download_file()
            content = await download.readall()

import asyncio
asyncio.run(datalake_operations())

Best Practices

  1. Pick sync OR async and stay consistent. Do not mix azure.storage.filedatalake sync clients with azure.storage.filedatalake.aio async clients in the same call path. Choose one mode per module.
  2. Always use context managers for clients and async credentials. Wrap every client in with DataLakeServiceClient(...) as client: (sync) or async with DataLakeServiceClient(...) as client: (async). For async DefaultAzureCredential from azure.identity.aio, also use async with credential: so tokens and transports are cleaned up.
  3. Use DefaultAzureCredential for portable auth across local dev and Azure (avoid connection strings / API keys when possible).
  4. Use hierarchical namespace for file system semantics
  5. Use append_data + flush_data for large file uploads
  6. Set ACLs at directory level and inherit to children
  7. Use async client for high-throughput scenarios
  8. Use get_paths with recursive=True for full directory listing
  9. Set metadata for custom file attributes
  10. Consider Blob API for simple object storage use cases

Reference Files

FileContents
references/capabilities.mdAdditional non-hero capabilities, operation-group coverage, and production checklists.
references/non-hero-scenarios.mdDedicated non-hero examples for secondary/advanced scenarios.

© microsoft, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py of microsoft/skills.

  • SKILL.md
  • references/capabilities.md
  • references/non-hero-scenarios.md

Open the folder on GitHubat commit 3898ec8

Compare with similar skills

Azure Storage File Datalake Py next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Azure Storage File Datalake Py compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Azure Storage File Datalake Py this skillmicrosoft/skills3.1k—~2kAutomated safety check: PassMIT
Azure Storage File Datalake Pyaiskillstore/marketplace4304 repos~1.5kAutomated safety check: PassNone
Azure Architecture Autopilotgithub/awesome-copilot40k1 repos~1.9kAutomated safety check: PassMIT
Terraform Azurerm Set Diff Analyzergithub/awesome-copilot40k1 repos~547Automated safety check: PassMIT
Osmo Lerobot Trainingmicrosoft/physical-ai-toolchain123—~3.8kAutomated safety check: NotesMIT
Azure AI Deploytimothywarner-org/claude-code224—~731Automated safety check: NotesMIT

Similar skills

  • Azure Storage File Datalake Py

    aiskillstore/marketplace

    Azure Data Lake Storage Gen2 SDK for Python. An agent skill from aiskillstore/marketplace.

    430 GitHub starsUsed in 4 repos~1.5k tokens
    DevOps & CloudAuto-check passed
  • Azure Architecture Autopilot

    github/awesome-copilot

    Official

    Designs Azure infrastructure from a natural-language description, or diagrams an existing resource group, then refines the design through conversation and deploys it with Bicep.

    40k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Official

    Analyze Terraform plan JSON output for AzureRM Provider to distinguish between false-positive diffs (order-only changes in Set-type attributes) and actual resource changes.

    40k GitHub starsUsed in 1 repo~547 tokens
    DevOps & CloudAuto-check passed
  • Osmo Lerobot Training

    microsoft/physical-ai-toolchain

    Official

    Submit, monitor, analyze, and evaluate LeRobot imitation learning training jobs on OSMO with Azure ML MLflow integration and inference evaluation - Brought to you by microsoft/physical-ai-toolchain

    123 GitHub stars~3.8k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Azure AI Deploy

    timothywarner-org/claude-code

    Ship a Python generative-AI app to Azure the keyless way, using DefaultAzureCredential and azd.

    224 GitHub stars~731 tokensUpdated 2 mo ago
    DevOps & CloudAuto-check: notes
  • Audit ML Pipeline

    probabl-ai/skills

    Read-only audit of one persisted skore report: audit/NN<stem.py (jupytext percent), 1:1 with experiments/ and journal/.

    138 GitHub stars~9.4k tokensUpdated yesterday
    DevOps & CloudAuto-check passed

More from microsoft/skills

All 150 skills in this repo
  • Official

    Covers producer, consumer, and checkpoint-store setup for Azure Event Hubs streaming in Python, with Entra ID auth and partition targeting.

    3.1k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Official

    Builds podcast-style audio narration from text with Azure OpenAI's GPT Realtime Mini over WebSocket, from a Python FastAPI backend to a React player.

    3.1k GitHub starsUsed in 1 repo~947 tokens
    Auto-check passed
  • Frontend UI Dark TS

    microsoft/skills

    Official

    Build dark-themed React applications using Tailwind CSS with custom theming, glassmorphism effects, and Framer Motion animations.

    3.1k GitHub starsUsed in 5 repos~3.6k tokens
    Auto-check passed
  • Pydantic Models Py

    microsoft/skills

    Official

    Create Pydantic models following the multi-model pattern with Base, Create, Update, Response, and InDB variants.

    3.1k GitHub starsUsed in 5 repos~496 tokens
    Auto-check passed
  • Official

    Reference for building on Microsoft Foundry with the azure-ai-projects Python SDK: project clients, versioned agents, evaluations, connections, datasets and indexes.

    3.1k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Skill Creator

    microsoft/skills

    Official

    Guide for creating effective skills for AI coding agents working with Azure SDKs and Microsoft Foundry services.

    3.1k GitHub starsUsed in 5 repos~17k tokens
    Auto-check passed

Questions about Azure Storage File Datalake Py

What does Azure Storage File Datalake Py do?

Azure Data Lake Storage Gen2 SDK for Python. An agent skill from microsoft/skills. Azure Storage File Datalake Py is an agent skill from microsoft/skills, published by the product's own GitHub organization. Azure Data Lake Storage Gen2 SDK for Python.

When should I use Azure Storage File Datalake Py?

Azure Storage File Datalake Py fits situations like: hierarchical file systems; big data analytics; file/directory operations.

How do I install Azure Storage File Datalake Py in Claude Code?

Run `npx skills add microsoft/skills --skill azure-storage-file-datalake-py -a claude-code`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py in microsoft/skills) into .claude/skills/azure-storage-file-datalake-py in your project. Claude Code loads it when a task matches its description.

How do I install Azure Storage File Datalake Py in Codex?

Run `npx skills add microsoft/skills --skill azure-storage-file-datalake-py -a codex`. Or copy the skill folder (.github/plugins/azure-sdk-python/skills/azure-storage-file-datalake-py in microsoft/skills) into .agents/skills/azure-storage-file-datalake-py in your project. Codex loads it when a task matches its description.

Can I use Azure Storage File Datalake Py in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add microsoft/skills --skill azure-storage-file-datalake-py -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/azure-storage-file-datalake-py, .gemini/skills/azure-storage-file-datalake-py, .github/skills/azure-storage-file-datalake-py and .opencode/skills/azure-storage-file-datalake-py in your project.

What does Azure Storage File Datalake Py need to run?

Going by SKILL.md and its folder, Azure Storage File Datalake Py needs the command-line tools its instructions call (pip) and credentials named AZURE_TOKEN_CREDENTIALS. Our summary lists: Python 3.

Does Azure Storage File Datalake Py access the network?

SKILL.md names 1 domain. In commands or code: learn.microsoft.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Azure Storage File Datalake Py safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Azure Storage File Datalake Py use?

Azure Storage File Datalake Py is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Azure Storage File Datalake Py use?

About 2k tokens (SKILL.md is roughly 8.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 947 tokens, read only when the agent opens those files.

What are the alternatives to Azure Storage File Datalake Py?

Skills that share tags, products or a category with Azure Storage File Datalake Py: Azure Storage File Datalake Py (aiskillstore/marketplace, 430 stars), Azure Architecture Autopilot (github/awesome-copilot, 40k stars), Terraform Azurerm Set Diff Analyzer (github/awesome-copilot, 40k stars) and Osmo Lerobot Training (microsoft/physical-ai-toolchain, 123 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Azure Storage File Datalake Py?

microsoft (a GitHub organization, an official publisher) maintains it in microsoft/skills, which has 3,094 GitHub stars. The repository holds 150 skills in this directory. The repository was last updated on October 9, 2026.

Source: microsoft/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.