如何在Azure Data Factory V2中通过PySpark活动连接并列出Blob存储容器文件
Hey Jess, great question! Let’s walk through exactly how to connect to an Azure Blob Storage container using a PySpark activity in Azure Data Factory (ADF) v2, and list all the files inside the container. I’ll break this down into actionable steps with code examples you can reuse.
Prerequisites
First, make sure you have these ready:
- An Azure Storage account with a Blob container (and files inside it, obviously!)
- An ADF v2 instance up and running
- A Spark compute resource linked to ADF (this can be an Azure Synapse Spark Pool, HDInsight Spark Cluster, or even a serverless Spark pool)
- Your storage account access key (or a SAS token) – we’ll cover secure ways to use this later
Step 1: Configure a Blob Storage Linked Service (Optional but Recommended)
While you can hardcode credentials in your PySpark script (not ideal!), using an ADF Linked Service keeps your sensitive info secure and manageable. Here’s how to set it up:
- In ADF Studio, go to Manage > Linked Services > New
- Search for Azure Blob Storage and select it
- Choose your authentication method (Account Key is simplest for testing; use SAS or Managed Identity for production)
- Fill in your storage account details and save the Linked Service
Step 2: Create a PySpark Activity in Your Pipeline
- Navigate to Author in ADF Studio and create a new pipeline
- Drag a PySpark activity from the Activities pane into your pipeline canvas
- Go to the Compute tab of the activity and select your linked Spark compute resource
- (Optional) If you want to pass dynamic values, go to the Parameters tab of your pipeline and add parameters like
storageAccountName,containerName, andstorageAccountKey(we’ll use these in the script)
Step 3: Write the PySpark Code to List Blob Files
Now for the fun part – the code. I’ll share two approaches: one using pipeline parameters (secure, no hardcoding) and a bonus recursive file listing option.
Basic File Listing (Root Directory)
This script lists all files and folders in the root of your Blob container:
# Pull parameters from your ADF pipeline (replace with your parameter names if different) storage_account_name = "@pipeline().parameters.storageAccountName" storage_account_key = "@pipeline().parameters.storageAccountKey" container_name = "@pipeline().parameters.containerName" # Configure Spark to access Blob Storage spark.conf.set( f"fs.azure.account.key.{storage_account_name}.blob.core.windows.net", storage_account_key ) # Use wasbs protocol for standard Blob Storage; use abfs if you're on ADLS Gen2 for better performance container_path = f"wasbs://{container_name}@{storage_account_name}.blob.core.windows.net/" # List items in the container root items = dbutils.fs.ls(container_path) # Print the results print(f"Listing items in container '{container_name}':") for item in items: item_type = "File" if item.isFile() else "Folder" print(f"- {item_type}: {item.name} | Size: {item.size} bytes | Path: {item.path}")
Recursive File Listing (All Subfolders)
If you need to list every file in the container (including nested subfolders), use this recursive function:
# Pull pipeline parameters storage_account_name = "@pipeline().parameters.storageAccountName" storage_account_key = "@pipeline().parameters.storageAccountKey" container_name = "@pipeline().parameters.containerName" # Configure Spark spark.conf.set( f"fs.azure.account.key.{storage_account_name}.blob.core.windows.net", storage_account_key ) container_path = f"wasbs://{container_name}@{storage_account_name}.blob.core.windows.net/" # Recursive function to list all files all_files = [] def scan_directory(path): for item in dbutils.fs.ls(path): if item.isFile(): all_files.append(item) else: scan_directory(item.path) # Run the scan scan_directory(container_path) # Output results print(f"Total files found: {len(all_files)}") for file in all_files: print(f"- File: {file.name} | Size: {file.size} bytes | Full Path: {file.path}")
Pro Security Tip
Instead of passing the storage key as a pipeline parameter, store it in Azure Key Vault. Then:
- Create a Key Vault Linked Service in ADF
- Use a Lookup Activity to retrieve the key from Key Vault
- Pass the retrieved key to your PySpark activity as a parameter
This way, your key is never exposed in pipeline definitions or scripts.
Step 4: Test and Validate
- Save your pipeline and click Debug to run it
- Once the PySpark activity completes, go to the Output tab of the activity to see the printed file list
- If you run into errors, check:
- Spark cluster network access (ensure it can reach your storage account – disable firewall temporarily for testing if needed)
- Correct storage account name/key/container name
- Proper protocol (wasbs vs abfs)
That’s it! You should now be able to connect to Blob Storage from ADF’s PySpark activity and list your files easily.
内容的提问来源于stack exchange,提问作者Jess

