如何让PySpark-Synapse Notebook读取无文件时不报错?
解决Synapse Notebook读取Parquet无匹配文件时的报错问题
方案1:手动指定Schema + 异常捕获
当没有匹配文件时,Spark无法自动推断Schema,提前定义好目标Schema,再通过捕获异常返回空DataFrame:
from pyspark.sql.types import StructType, StructField, StringType, IntegerType # 按实际数据类型调整 # 替换为你的Parquet文件真实结构 custom_schema = StructType([ StructField("col1", StringType(), nullable=True), StructField("col2", IntegerType(), nullable=True) ]) try: ReadDF = spark.read.load(readPath, format="parquet", schema=custom_schema, modifiedBefore=PLP___EndDate, modifiedAfter=PLP___StartDate) except Exception as e: # 捕获无文件匹配的异常,返回空DataFrame ReadDF = spark.createDataFrame([], schema=custom_schema)
方案2:先检查文件再读取
通过Hadoop文件系统API提前校验指定路径下是否有符合时间条件的文件,再决定是否执行读取:
from py4j.java_gateway import java_import java_import(spark._jvm, "org.apache.hadoop.fs.Path") java_import(spark._jvm, "org.apache.hadoop.fs.FileSystem") fs = spark._jvm.FileSystem.get(spark._jsc.hadoopConfiguration()) path = spark._jvm.Path(readPath) # 遍历文件,校验修改时间是否在目标范围内 file_statuses = fs.listStatus(path) has_matching_files = any( stat.getModificationTime() >= PLP___StartDate.timestamp() * 1000 and stat.getModificationTime() <= PLP___EndDate.timestamp() * 1000 for stat in file_statuses ) if has_matching_files: ReadDF = spark.read.load(readPath, format="parquet", modifiedBefore=PLP___EndDate, modifiedAfter=PLP___StartDate) else: # 无匹配文件时返回空DataFrame,复用之前定义的custom_schema ReadDF = spark.createDataFrame([], schema=custom_schema)
方案3:开启Spark忽略缺失文件配置
Spark提供ignoreMissingFiles参数,开启后会忽略不存在的文件,结合手动指定Schema即可处理无文件场景:
spark.conf.set("spark.sql.files.ignoreMissingFiles", "true") # 手动指定Schema避免推断失败 ReadDF = spark.read.load(readPath, format="parquet", schema=custom_schema, modifiedBefore=PLP___EndDate, modifiedAfter=PLP___StartDate) # 无匹配文件时,ReadDF会直接返回空表,不会抛出异常
内容的提问来源于stack exchange,提问作者Oblivi0n
相关产品推荐
相关产品推荐

