Apache Spark集群临时文件存储路径与清理方案问询
Hey there! Since you're setting up a 10-node Spark cluster, let's walk through all your questions about temporary files clearly—they're a common pain point, but easy to manage once you know the ins and outs.
1. Where Are Spark's Temporary Files Stored?
First off: all temporary files generated during job execution (like shuffle intermediate data, persisted RDDs, etc.) live on the corresponding Worker nodes, not the Master. The Master's job is just scheduling resources and managing job submissions—it never stores runtime temporary data. That makes sense, right? You don't want the Master becoming a bottleneck with all that file I/O.
2. Default Temporary Directory Paths & How to Configure Them
Default Paths
By default, Spark uses two main locations for temp files:
- System temp directory: On Linux/Unix, this is
/tmp(or whatever thejava.io.tmpdirsystem variable points to). - Worker workspace: Shuffle-specific temp files also get stored in subfolders under
${SPARK_HOME}/workon each Worker node.
Customizing the Paths
You can tweak these locations in a few ways:
- Via
spark-defaults.conf(persistent cluster-wide setting)
Add these lines to set a custom temp directory (you can list multiple paths separated by commas for load balancing):spark.local.dir /path/to/your/custom/temp/folder,/another/path # Optional: Separate path for shuffle files spark.shuffle.service.dir /path/to/shuffle-specific/temp - Per-job via
spark-submit
Override the temp dir when submitting a single job:spark-submit --conf spark.local.dir=/path/to/custom/temp your_job.jar - Worker root workspace via
spark-env.sh
Set the base directory for Worker operations (temp files will live in subfolders here):export SPARK_WORKER_DIR=/path/to/worker/workspace
3. Handling Full Temp Directories & Automatic Cleanup
Manual & Job-Level Cleanup
If your temp dir fills up, here are immediate fixes:
- Unpersist unused RDDs: In your job code, call
rdd.unpersist()on any persisted RDDs you no longer need—this frees up both memory and disk temp data. - Adjust shuffle file retention: By default, shuffle files stick around for 2 hours (7200 seconds) after a job finishes. Shorten this with these configs:
spark.shuffle.service.cleanup.interval 600 # Clean every 10 minutes spark.shuffle.service.delete.delay 60 # Delete shuffle files 1 minute after job completion
Does spark.worker.cleanup.enabled=true Work?
Absolutely—this is the go-to config for automatic cleanup of old job data on Workers. But you'll want to pair it with related settings to make it effective:
Add these to spark-defaults.conf:
spark.worker.cleanup.enabled true spark.worker.cleanup.interval 3600 # Run cleanup every hour (default) spark.worker.cleanup.appDataTtl 86400 # Keep job data for 1 day (default)—older data gets deleted
Important note: This only cleans up data from completed/failed jobs, not currently running ones. So it won't interfere with active workloads, just clear out old clutter taking up space.
内容的提问来源于stack exchange,提问作者LennyMarshal

