MacOS下Spark写入文件报ParentNotDirectoryException问题求助
问题排查与解决
报错原因
Spark的DataFrameWriter.csv()方法要求传入的输出路径必须是目录路径,而非具体文件名。你当前传入的output_filepath是file:/path/to/my/data_file.tsv.gz,Spark会将这个路径视为要创建的父目录,而它显然不是一个目录,因此抛出ParentNotDirectoryException。
解决步骤
1. 拆分输出路径
将目标文件名与所在目录分离,例如:
- 输出目录:
/path/to/my/ - 目标文件名:
data_file.tsv.gz
2. 修改代码写入目录,再重命名文件
因为你用了coalesce(1),Spark会在指定目录下生成单个part文件,之后手动将其重命名为目标文件名即可。示例代码如下:
import os import shutil # 拆分路径 output_dir = "/path/to/my/" target_filename = "data_file.tsv.gz" # 写入到目录 spark_dataframe.coalesce(1).write.csv( output_dir, mode=mode, compression=compression, sep="\t" if tsv else ",", header=False, escape='"', ) # 找到生成的part文件并重命名 for filename in os.listdir(output_dir): if filename.startswith("part-") and filename.endswith(".gz"): part_file_path = os.path.join(output_dir, filename) target_file_path = os.path.join(output_dir, target_filename) # 如果目标文件已存在,先删除(根据mode参数调整) if os.path.exists(target_file_path): os.remove(target_file_path) shutil.move(part_file_path, target_file_path) # 删除Spark生成的_SUCCESS文件(可选) success_file = os.path.join(output_dir, "_SUCCESS") if os.path.exists(success_file): os.remove(success_file) break
补充说明
你提到相同代码在Ubuntu上运行正常,大概率是因为不同系统下Hadoop文件系统对路径的解析逻辑存在差异,但Spark的官方规范始终是写入目录而非直接写入文件名,因此统一按目录写入后重命名的方式更可靠。
内容的提问来源于stack exchange,提问作者led0764
相关产品推荐
相关产品推荐

