Spark读取含特殊字符(重音)文件名的文件失败问题咨询
解决Spark读取含特殊/重音字符文件名的XML文件问题
问题背景
在复现Databricks workshop的XML解析功能时,执行以下命令将样本数据解压至DBFS路径:
mkdir -p /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass wget https://synthetichealth.github.io/synthea-sample-data/downloads/synthea_sample_data_ccda_sep2019.zip -O ./synthea_sample_data_ccda_sep2019.zip unzip ./synthea_sample_data_ccda_sep2019.zip -d /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/
使用Spark 3.2.1(Databricks DBR 10.4 LTS)读取XML文件时,遇到FileNotFoundException,明明存在的文件(如Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml)无法被找到。读取代码如下:
spark.conf.set("spark.sql.caseSensitive", "true") df = ( spark.read.format('xml') .option("rowTag", "ClinicalDocument") .load('/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/') )
报错信息:
org.apache.spark.SparkException: Job aborted due to stage failure: Task 43 in stage 251.0 failed 4 times, most recent failure: Lost task 43.3 in stage 251.0 (TID 494) (10.139.64.6 executor 0): java.io.FileNotFoundException: /4842022074360943/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml
问题原因
核心原因是解压时的编码不匹配:unzip命令在默认情况下可能使用非UTF-8编码(如CP437)处理带重音的文件名,导致DBFS中存储的文件名实际是乱码,而Spark读取时使用UTF-8解析路径,因此无法匹配到正确的文件。
解决方案
1. 重新解压时指定UTF-8编码
删除原有解压的文件,重新执行解压命令并添加-O UTF-8参数,强制用UTF-8编码解析文件名:
# 删除原有目录 rm -rf /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ # 重新创建目录 mkdir -p /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass # 重新下载(如果已下载可跳过) wget https://synthetichealth.github.io/synthea-sample-data/downloads/synthea_sample_data_ccda_sep2019.zip -O ./synthea_sample_data_ccda_sep2019.zip # 指定UTF-8编码解压 unzip -O UTF-8 ./synthea_sample_data_ccda_sep2019.zip -d /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/
2. 验证DBFS中的文件名
使用Databricks的dbutils工具查看目标目录下的文件名,确认重音字符正常显示:
display(dbutils.fs.ls("/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/"))
如果文件名显示正常(如Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml),再重新执行Spark读取代码即可。
3. 配置Spark路径编码(可选)
如果仍有问题,可以设置Spark的路径过滤器编码为UTF-8,确保路径解析时使用正确编码:
spark.conf.set("spark.files.pathFilterEncoding", "UTF-8") spark.conf.set("spark.sql.caseSensitive", "true") df = ( spark.read.format('xml') .option("rowTag", "ClinicalDocument") .load('/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/') )
4. 批量重命名文件(备选方案)
如果解压后的文件名确实存在编码问题且无法通过重新解压解决,可以批量重命名文件,移除或替换特殊字符。示例代码:
import os from pathlib import Path dbfs_path = "/dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/" for file in Path(dbfs_path).iterdir(): if file.is_file(): # 替换重音字符为非重音版本,或直接移除特殊字符 new_name = file.name.replace("ñ", "n") # 示例替换ñ为n os.rename(file, os.path.join(dbfs_path, new_name))
执行后再用Spark读取重命名后的文件。
内容的提问来源于stack exchange,提问作者xuxu
相关产品推荐
相关产品推荐

