You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpark保存DataFrame时如何避免文件名中的随机串?

How to Generate PySpark CSV Output Files with Standard Part Names (No Random Strings)

Great question! Those random string suffixes in your output filenames are temporary artifacts from Spark's default file-writing mechanism. Let me break down why they appear and how to fix this:

Why the Random Strings Exist

By default, Spark uses version 2 of the Hadoop FileOutputCommitter algorithm. This approach first writes task outputs to temporary directories with randomized filenames, then renames them to the standard part-xxxxx format only after all tasks complete successfully. If you check the output directory before the job finishes (or if there's a partial failure), you'll see those messy random-named files instead of the clean part files you expect.

Fix 1: Switch to the Legacy Output Committer Algorithm

To get clean part-00001.csv.gz style filenames directly, switch to version 1 of the FileOutputCommitter algorithm. This writes files straight to the final directory with standard names, skipping the temporary file step entirely.

You can set this either globally in your SparkSession:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("StandardPartNames") \
    .config("spark.hadoop.mapreduce.fileoutputcommitter.algorithm.version", "1") \
    .getOrCreate()

Or apply it only to your specific write operation (no global config changes needed):

df.write.format("csv") \
    .options(header='false', inferschema='true', sep="|") \
    .option("codec", "org.apache.hadoop.io.compress.GzipCodec") \
    .option("mapreduce.fileoutputcommitter.algorithm.version", "1") \
    .save("path")

Fix 2: Control the Number of Output Files (Optional)

If you want to specify exactly how many part files are generated (instead of letting Spark decide based on partitions), use coalesce() or repartition() to adjust your DataFrame's partition count before writing:

# Generate 3 part files (adjust the number to match your needs)
df.coalesce(3) \
    .write.format("csv") \
    .options(header='false', inferschema='true', sep="|") \
    .option("codec", "org.apache.hadoop.io.compress.GzipCodec") \
    .option("mapreduce.fileoutputcommitter.algorithm.version", "1") \
    .save("path")

This will give you part-00000.csv.gz, part-00001.csv.gz, part-00002.csv.gz — no random strings in sight!

A Quick Note on Tradeoffs

Version 1 of the committer algorithm is simpler and gives you clean filenames, but it has a minor downside: in highly distributed environments, there's a tiny risk of file write conflicts (though Spark's partitioning logic usually prevents this). For most use cases, this is a negligible tradeoff compared to getting the filename format you want.

内容的提问来源于stack exchange,提问作者Sai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 03:57:17