将Azure Data Lake数据加载到DSVM的Jupyter Notebook遇问题求助
Let's tackle your PySpark problem first since that's what you're focused on, and we'll cover the Python SDK snag too.
PySpark Error: No FileSystem for scheme: adl
This error pops up because Spark's Hadoop environment doesn't have the right drivers or configuration to interact with Azure Data Lake Storage (ADLS) Gen1. Here's how to fix it step by step:
1. Verify ADLS dependencies are present
DSVMs usually come with these pre-installed, but it's worth checking:
Run this in your Jupyter notebook to confirm the required JARs are in Spark's classpath:
sc.listJars().filter(lambda jar: 'datalake' in jar.lower()).collect()
If you get an empty list, you'll need to add the jars manually. Look for hadoop-azure-datalake.jar and azure-datalake-store-sdk.jar in /usr/hdp/current/hadoop-client/lib/ (a common DSVM path) and add them to Spark's classpath via the spark.driver.extraClassPath config.
2. Configure Spark for ADLS
Add these lines at the start of your PySpark script to set up the ADLS file system and authentication:
# Set ADLS file system implementation classes sc._jsc.hadoopConfiguration().set("fs.adl.impl", "org.apache.hadoop.fs.adl.AdlFileSystem") sc._jsc.hadoopConfiguration().set("fs.AbstractFileSystem.adl.impl", "org.apache.hadoop.fs.adl.Adl") # For interactive authentication (great for DSVMs since you're already signed in) from pyspark.sql import SparkSession spark = SparkSession.builder.appName("ADLS-Load").getOrCreate() # If you prefer using a Service Principal (better for automated jobs), uncomment and fill in these values: # sc._jsc.hadoopConfiguration().set("dfs.adls.oauth2.access.token.provider.type", "ClientCredential") # sc._jsc.hadoopConfiguration().set("dfs.adls.oauth2.client.id", "<YOUR-CLIENT-ID>") # sc._jsc.hadoopConfiguration().set("dfs.adls.oauth2.credential", "<YOUR-CLIENT-SECRET>") # sc._jsc.hadoopConfiguration().set("dfs.adls.oauth2.refresh.url", "https://login.microsoftonline.com/<YOUR-TENANT-ID>/oauth2/token")
3. Retry loading your data
Once the config is set, run your original code again—it should now connect properly:
hvacText = sc.textFile("adl://name.azuredatalakestore.net/file_to_read.csv") hvacText.count()
Python SDK Error: DatalakeRESTException: LISTSTATUS
This is almost always a permission or token scope issue, even though you're the owner. Try these fixes:
- Specify the ADLS resource ID when authenticating to ensure your token has the right permissions:
from azure.datalake.store import core, lib, multithread token = lib.auth(resource='https://datalake.azure.net/') adl = core.AzureDLFileSystem(token, store_name='store_name') adl.ls('/')
- Double-check that your account has the Data Lake Storage Gen1 Contributor role (or higher) on the storage account. Sometimes owner permissions take a few minutes to propagate across Azure, so wait a bit if you just assigned the role.
内容的提问来源于stack exchange,提问作者TrnKh

