You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在PySpark读取ADLS中带通配符的文件到DataFrame前检查存在性?

Solution: Check for Matching Files Before Reading in PySpark Notebook

Got it, let's solve this problem step by step. You need to verify if there are files matching your wildcard pattern in the ADLS directory before loading them into a PySpark DataFrame, and halt the Notebook immediately if no files exist. Here are two reliable approaches depending on your environment:

Approach 1: Databricks Notebook (Most Common for ADLS Mounts)

Since your path uses /mnt/, I assume you're working in a Databricks environment with ADLS mounted. We'll use dbutils to check for matching files and terminate the Notebook if none are found:

# Define your path and wildcard pattern
kaka_adls_path = "/mnt/kaka/pre/Source_Files/"
kaka_fac_address = "*FacKaka*.txt"
full_pattern = f"{kaka_adls_path}/{kaka_fac_address}"

# List all files matching the wildcard pattern
matching_files = dbutils.fs.ls(full_pattern)

# Check if any files exist
if not matching_files:
    print("❌ No files matching the pattern found. Terminating Notebook execution.")
    # This command immediately stops the Notebook
    dbutils.notebook.exit("Terminated: No matching files detected")
else:
    # Read the files into a PySpark DataFrame
    kaka_fac_address_df = spark.read.text(full_pattern)
    print(f"✅ Successfully read {len(matching_files)} file(s) into DataFrame.")

Key Notes:

  • dbutils.fs.ls(full_pattern) automatically handles wildcards and returns a list of file objects (each with path, name, and size attributes).
  • dbutils.notebook.exit() is the cleanest way to stop a Databricks Notebook—it exits immediately without running any subsequent code.

Approach 2: Generic PySpark (Non-Databricks Environments)

If you're working in a standard PySpark environment (not Databricks), use Hadoop's FileSystem API to check for matching files:

from pyspark.sql import SparkSession

# Initialize Spark Session (if not already done)
spark = SparkSession.builder.appName("FileCheck").getOrCreate()
sc = spark.sparkContext

# Define path and pattern
kaka_adls_path = "/mnt/kaka/pre/Source_Files/"
kaka_fac_address = "*FacKaka*.txt"
full_pattern = f"{kaka_adls_path}/{kaka_fac_address}"

# Access Hadoop FileSystem
hadoop_conf = sc._jsc.hadoopConfiguration()
path = sc._jvm.org.apache.hadoop.fs.Path(full_pattern)
fs = path.getFileSystem(hadoop_conf)

# Check for matching files
matching_files = fs.globStatus(path)

if not matching_files:
    print("❌ No files matching the pattern found. Terminating execution.")
    # Raise an exception to halt execution
    raise Exception("Execution stopped: No matching files found")
else:
    kaka_fac_address_df = spark.read.text(full_pattern)
    print(f"✅ Successfully read {len(matching_files)} file(s) into DataFrame.")

Key Notes:

  • fs.globStatus(path) returns an array of file status objects for all files matching the wildcard pattern.
  • Raising an Exception ensures the script stops immediately—adjust error handling as needed for your workflow.

内容的提问来源于stack exchange,提问作者Anonymous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 10:31:14