You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark读取含特殊字符(重音)文件名的文件失败问题咨询

解决Spark读取含特殊/重音字符文件名的XML文件问题

问题背景

在复现Databricks workshop的XML解析功能时,执行以下命令将样本数据解压至DBFS路径:

mkdir -p /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass
wget https://synthetichealth.github.io/synthea-sample-data/downloads/synthea_sample_data_ccda_sep2019.zip -O ./synthea_sample_data_ccda_sep2019.zip 
unzip ./synthea_sample_data_ccda_sep2019.zip -d /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/

使用Spark 3.2.1(Databricks DBR 10.4 LTS)读取XML文件时,遇到FileNotFoundException,明明存在的文件(如Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml)无法被找到。读取代码如下:

spark.conf.set("spark.sql.caseSensitive", "true")
df = (
  spark.read.format('xml')
   .option("rowTag", "ClinicalDocument")
   .load('/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/')
)

报错信息:

org.apache.spark.SparkException: Job aborted due to stage failure: Task 43 in stage 251.0 failed 4 times, most recent failure: Lost task 43.3 in stage 251.0 (TID 494) (10.139.64.6 executor 0): java.io.FileNotFoundException: /4842022074360943/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml

问题原因

核心原因是解压时的编码不匹配:unzip命令在默认情况下可能使用非UTF-8编码(如CP437)处理带重音的文件名,导致DBFS中存储的文件名实际是乱码,而Spark读取时使用UTF-8解析路径,因此无法匹配到正确的文件。

解决方案

1. 重新解压时指定UTF-8编码

删除原有解压的文件,重新执行解压命令并添加-O UTF-8参数,强制用UTF-8编码解析文件名:

# 删除原有目录
rm -rf /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/
# 重新创建目录
mkdir -p /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass
# 重新下载(如果已下载可跳过)
wget https://synthetichealth.github.io/synthea-sample-data/downloads/synthea_sample_data_ccda_sep2019.zip -O ./synthea_sample_data_ccda_sep2019.zip 
# 指定UTF-8编码解压
unzip -O UTF-8 ./synthea_sample_data_ccda_sep2019.zip -d /dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/

2. 验证DBFS中的文件名

使用Databricks的dbutils工具查看目标目录下的文件名,确认重音字符正常显示:

display(dbutils.fs.ls("/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/"))

如果文件名显示正常(如Amalia471_Magaña874_3912696a-0aef-492e-83ef-468262b82966.xml),再重新执行Spark读取代码即可。

3. 配置Spark路径编码(可选)

如果仍有问题,可以设置Spark的路径过滤器编码为UTF-8,确保路径解析时使用正确编码:

spark.conf.set("spark.files.pathFilterEncoding", "UTF-8")
spark.conf.set("spark.sql.caseSensitive", "true")
df = (
  spark.read.format('xml')
   .option("rowTag", "ClinicalDocument")
   .load('/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/')
)

4. 批量重命名文件(备选方案)

如果解压后的文件名确实存在编码问题且无法通过重新解压解决,可以批量重命名文件,移除或替换特殊字符。示例代码:

import os
from pathlib import Path

dbfs_path = "/dbfs/user/hive/warehouse/hls_cms_source.db/raw_files/synthea_mass/ccda/"
for file in Path(dbfs_path).iterdir():
    if file.is_file():
        # 替换重音字符为非重音版本,或直接移除特殊字符
        new_name = file.name.replace("ñ", "n")  # 示例替换ñ为n
        os.rename(file, os.path.join(dbfs_path, new_name))

执行后再用Spark读取重命名后的文件。

内容的提问来源于stack exchange,提问作者xuxu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 12:20:36