如何在PySpark读取ADLS中带通配符的文件到DataFrame前检查存在性?
Solution: Check for Matching Files Before Reading in PySpark Notebook
Got it, let's solve this problem step by step. You need to verify if there are files matching your wildcard pattern in the ADLS directory before loading them into a PySpark DataFrame, and halt the Notebook immediately if no files exist. Here are two reliable approaches depending on your environment:
Approach 1: Databricks Notebook (Most Common for ADLS Mounts)
Since your path uses /mnt/, I assume you're working in a Databricks environment with ADLS mounted. We'll use dbutils to check for matching files and terminate the Notebook if none are found:
# Define your path and wildcard pattern kaka_adls_path = "/mnt/kaka/pre/Source_Files/" kaka_fac_address = "*FacKaka*.txt" full_pattern = f"{kaka_adls_path}/{kaka_fac_address}" # List all files matching the wildcard pattern matching_files = dbutils.fs.ls(full_pattern) # Check if any files exist if not matching_files: print("❌ No files matching the pattern found. Terminating Notebook execution.") # This command immediately stops the Notebook dbutils.notebook.exit("Terminated: No matching files detected") else: # Read the files into a PySpark DataFrame kaka_fac_address_df = spark.read.text(full_pattern) print(f"✅ Successfully read {len(matching_files)} file(s) into DataFrame.")
Key Notes:
dbutils.fs.ls(full_pattern)automatically handles wildcards and returns a list of file objects (each withpath,name, andsizeattributes).dbutils.notebook.exit()is the cleanest way to stop a Databricks Notebook—it exits immediately without running any subsequent code.
Approach 2: Generic PySpark (Non-Databricks Environments)
If you're working in a standard PySpark environment (not Databricks), use Hadoop's FileSystem API to check for matching files:
from pyspark.sql import SparkSession # Initialize Spark Session (if not already done) spark = SparkSession.builder.appName("FileCheck").getOrCreate() sc = spark.sparkContext # Define path and pattern kaka_adls_path = "/mnt/kaka/pre/Source_Files/" kaka_fac_address = "*FacKaka*.txt" full_pattern = f"{kaka_adls_path}/{kaka_fac_address}" # Access Hadoop FileSystem hadoop_conf = sc._jsc.hadoopConfiguration() path = sc._jvm.org.apache.hadoop.fs.Path(full_pattern) fs = path.getFileSystem(hadoop_conf) # Check for matching files matching_files = fs.globStatus(path) if not matching_files: print("❌ No files matching the pattern found. Terminating execution.") # Raise an exception to halt execution raise Exception("Execution stopped: No matching files found") else: kaka_fac_address_df = spark.read.text(full_pattern) print(f"✅ Successfully read {len(matching_files)} file(s) into DataFrame.")
Key Notes:
fs.globStatus(path)returns an array of file status objects for all files matching the wildcard pattern.- Raising an
Exceptionensures the script stops immediately—adjust error handling as needed for your workflow.
内容的提问来源于stack exchange,提问作者Anonymous
相关产品推荐
相关产品推荐

