You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在指定路径的新文件夹写入CSV(解决FileNotFoundError)

解决Pandas写入CSV时的FileNotFoundError及Spark实现方案

Pandas问题解决

错误原因

  1. 路径拼接不规范:手动拼接路径时,依赖配置中路径末尾的斜杠保证格式正确,一旦配置变更就会出现路径错误;且目标父文件夹unmatched_column_data不存在,Pandas的to_csv方法无法自动创建不存在的目录。
  2. 未提前创建目标目录:Pandas仅能向已存在的目录写入文件,不会自动生成缺失的父文件夹。

修复代码

导入os模块,用系统原生方法处理路径拼接与目录创建,避免手动拼接的潜在问题:

import os

# 提取配置中的路径参数
config = obj['source_and_destination_details']
base_dir = config['local_uri_write']
target_folder = config['folder_name_for_unmatched_column_data']
file_name = config['file_name_for_unmatched_column_data']

# 拼接完整文件路径
full_path = os.path.join(base_dir, target_folder, file_name)

# 创建目标文件夹(已存在则跳过)
os.makedirs(os.path.dirname(full_path), exist_ok=True)

# 写入CSV文件
df.reindex(idx).to_csv(full_path, index=False)

Spark实现方式

Spark写入CSV的逻辑与Pandas不同:Spark会将指定路径作为文件夹,在其中生成多个分区文件(part-*.csv)、_SUCCESS标识文件等,这是分布式处理的默认行为。

基础分布式写入(推荐大数据场景)

# 假设已有Spark DataFrame:spark_df
output_dir = os.path.join(config['local_uri_write'], config['folder_name_for_unmatched_column_data'])

spark_df.write \
    .mode("overwrite")  # 可选模式:append/ignore/error(默认)
    .option("header", "true")  # 写入表头
    .csv(output_dir)

执行后,output_dir目录下会生成多个分区CSV文件,适合大数据量的分布式存储需求。

生成单个CSV文件(小数据场景适用)

如果业务要求生成单个文件,可通过coalesce(1)合并分区(注意:大数据量下不建议使用,会将数据集中到单个节点,影响性能):

spark_df.coalesce(1) \
    .write \
    .mode("overwrite") \
    .option("header", "true") \
    .csv(output_dir)

执行后,output_dir下会生成一个part-*.csv文件,可手动或通过脚本重命名为demo.csv。

内容的提问来源于stack exchange,提问作者Preeti Maddi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 17:15:10