You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将本地Spark实例连接至Azure Blob存储?

Connecting Local Spark to Azure Blob Storage: A Step-by-Step Guide

Absolutely! You can absolutely connect a local Spark instance to Azure Blob Storage—this is a common scenario, even if official docs and online resources often prioritize Databricks setups. Let’s break down exactly how to make this work, with practical examples and key considerations.

1. Ensure You Have the Right Dependencies

Spark relies on Hadoop’s Azure integration libraries to interact with Blob Storage. You’ll need two main packages, and their versions must match the Hadoop version bundled with your Spark distribution (check via spark-submit --version):

  • org.apache.hadoop:hadoop-azure: Handles Hadoop’s filesystem interface for Azure
  • com.microsoft.azure:azure-storage: Provides underlying Azure Storage SDK support

Option A: Use spark-submit with --packages

When launching your Spark job, include the packages directly:

spark-submit --packages org.apache.hadoop:hadoop-azure:3.3.4,com.microsoft.azure:azure-storage:8.6.5 your_spark_script.py

Adjust versions to match your Spark’s Hadoop version (e.g., if Spark uses Hadoop 2.7.x, use hadoop-azure:2.7.7 and azure-storage:7.0.0)

Option B: Add JARs to Spark’s Local Jars Directory

If you don’t want to specify packages every time, download the JARs and place them in your Spark installation’s jars folder (e.g., $SPARK_HOME/jars).

2. Configure Authentication

You have three primary authentication methods—pick the one that fits your use case:

Method 1: Storage Account Key (Simplest for Testing)

Use your Azure Storage account’s access key to authenticate. Add these configs to your SparkSession:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("LocalSparkToBlob") \
    .config("fs.azure.account.key.your-storage-account.blob.core.windows.net", "your-account-access-key") \
    .getOrCreate()

Method 2: SAS Token (More Secure, Limited Permissions)

A Shared Access Signature (SAS) token lets you grant restricted access (e.g., read-only, time-limited). Use it like this:

spark = SparkSession.builder \
    .appName("LocalSparkToBlobWithSAS") \
    .config("fs.azure.sas.your-container.your-storage-account.blob.core.windows.net", "your-sas-token-without-leading-?") \
    .getOrCreate()

Note: Omit the ? at the start of your SAS token when adding it to the config.

Method 3: Azure AD Service Principal (Production-Grade)

For enterprise environments, use an Azure AD service principal for role-based access control. You’ll need the client ID, client secret, and tenant ID:

spark = SparkSession.builder \
    .appName("LocalSparkToBlobWithAAD") \
    .config("fs.azure.account.auth.type.your-storage-account.blob.core.windows.net", "OAuth") \
    .config("fs.azure.account.oauth.provider.type.your-storage-account.blob.core.windows.net", "org.apache.hadoop.fs.azurebfs.oauth2.ClientCredsTokenProvider") \
    .config("fs.azure.account.oauth2.client.id.your-storage-account.blob.core.windows.net", "your-service-principal-client-id") \
    .config("fs.azure.account.oauth2.client.secret.your-storage-account.blob.core.windows.net", "your-service-principal-client-secret") \
    .config("fs.azure.account.oauth2.client.endpoint.your-storage-account.blob.core.windows.net", "https://login.microsoftonline.com/your-tenant-id/oauth2/token") \
    .getOrCreate()

3. Read/Write Data from Azure Blob

Once your SparkSession is configured, use the standard wasbs:// URI format to access Blob Storage:

  • Format: wasbs://<container-name>@<storage-account-name>.blob.core.windows.net/<file-path>

Example: Read a CSV File

df = spark.read.csv("wasbs://my-container@mystorageaccount.blob.core.windows.net/data/sample.csv", header=True, inferSchema=True)
df.show()

Example: Write Data to Blob

df.write.mode("overwrite").parquet("wasbs://my-container@mystorageaccount.blob.core.windows.net/output/processed_data")

Key Considerations

  • Network Access: Ensure your local machine can reach Azure Blob Storage (port 443 must be open; no corporate firewall blocks).
  • Version Compatibility: Mismatched Hadoop/Azure library versions are the #1 cause of errors—double-check versions match your Spark setup.
  • Security: Avoid hardcoding secrets (account keys, SAS tokens) in scripts. Use environment variables or secret management tools instead.
  • ADLS Gen2: If you’re using Azure Data Lake Storage Gen2, replace wasbs:// with abfss:// for better performance and additional features.

内容的提问来源于stack exchange,提问作者pratikpatre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 15:14:06