PySpark保存DataFrame时如何避免文件名中的随机串?
Great question! Those random string suffixes in your output filenames are temporary artifacts from Spark's default file-writing mechanism. Let me break down why they appear and how to fix this:
Why the Random Strings Exist
By default, Spark uses version 2 of the Hadoop FileOutputCommitter algorithm. This approach first writes task outputs to temporary directories with randomized filenames, then renames them to the standard part-xxxxx format only after all tasks complete successfully. If you check the output directory before the job finishes (or if there's a partial failure), you'll see those messy random-named files instead of the clean part files you expect.
Fix 1: Switch to the Legacy Output Committer Algorithm
To get clean part-00001.csv.gz style filenames directly, switch to version 1 of the FileOutputCommitter algorithm. This writes files straight to the final directory with standard names, skipping the temporary file step entirely.
You can set this either globally in your SparkSession:
from pyspark.sql import SparkSession spark = SparkSession.builder \ .appName("StandardPartNames") \ .config("spark.hadoop.mapreduce.fileoutputcommitter.algorithm.version", "1") \ .getOrCreate()
Or apply it only to your specific write operation (no global config changes needed):
df.write.format("csv") \ .options(header='false', inferschema='true', sep="|") \ .option("codec", "org.apache.hadoop.io.compress.GzipCodec") \ .option("mapreduce.fileoutputcommitter.algorithm.version", "1") \ .save("path")
Fix 2: Control the Number of Output Files (Optional)
If you want to specify exactly how many part files are generated (instead of letting Spark decide based on partitions), use coalesce() or repartition() to adjust your DataFrame's partition count before writing:
# Generate 3 part files (adjust the number to match your needs) df.coalesce(3) \ .write.format("csv") \ .options(header='false', inferschema='true', sep="|") \ .option("codec", "org.apache.hadoop.io.compress.GzipCodec") \ .option("mapreduce.fileoutputcommitter.algorithm.version", "1") \ .save("path")
This will give you part-00000.csv.gz, part-00001.csv.gz, part-00002.csv.gz — no random strings in sight!
A Quick Note on Tradeoffs
Version 1 of the committer algorithm is simpler and gives you clean filenames, but it has a minor downside: in highly distributed environments, there's a tiny risk of file write conflicts (though Spark's partitioning logic usually prevents this). For most use cases, this is a negligible tradeoff compared to getting the filename format you want.
内容的提问来源于stack exchange,提问作者Sai

