You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Apache Spark集群临时文件存储路径与清理方案问询

Hey there! Since you're setting up a 10-node Spark cluster, let's walk through all your questions about temporary files clearly—they're a common pain point, but easy to manage once you know the ins and outs.

Spark Temporary Files: Storage, Configuration, and Cleanup

1. Where Are Spark's Temporary Files Stored?

First off: all temporary files generated during job execution (like shuffle intermediate data, persisted RDDs, etc.) live on the corresponding Worker nodes, not the Master. The Master's job is just scheduling resources and managing job submissions—it never stores runtime temporary data. That makes sense, right? You don't want the Master becoming a bottleneck with all that file I/O.

2. Default Temporary Directory Paths & How to Configure Them

Default Paths

By default, Spark uses two main locations for temp files:

  • System temp directory: On Linux/Unix, this is /tmp (or whatever the java.io.tmpdir system variable points to).
  • Worker workspace: Shuffle-specific temp files also get stored in subfolders under ${SPARK_HOME}/work on each Worker node.

Customizing the Paths

You can tweak these locations in a few ways:

  • Via spark-defaults.conf (persistent cluster-wide setting)
    Add these lines to set a custom temp directory (you can list multiple paths separated by commas for load balancing):
    spark.local.dir /path/to/your/custom/temp/folder,/another/path
    # Optional: Separate path for shuffle files
    spark.shuffle.service.dir /path/to/shuffle-specific/temp
    
  • Per-job via spark-submit
    Override the temp dir when submitting a single job:
    spark-submit --conf spark.local.dir=/path/to/custom/temp your_job.jar
    
  • Worker root workspace via spark-env.sh
    Set the base directory for Worker operations (temp files will live in subfolders here):
    export SPARK_WORKER_DIR=/path/to/worker/workspace
    

3. Handling Full Temp Directories & Automatic Cleanup

Manual & Job-Level Cleanup

If your temp dir fills up, here are immediate fixes:

  • Unpersist unused RDDs: In your job code, call rdd.unpersist() on any persisted RDDs you no longer need—this frees up both memory and disk temp data.
  • Adjust shuffle file retention: By default, shuffle files stick around for 2 hours (7200 seconds) after a job finishes. Shorten this with these configs:
    spark.shuffle.service.cleanup.interval 600  # Clean every 10 minutes
    spark.shuffle.service.delete.delay 60       # Delete shuffle files 1 minute after job completion
    

Does spark.worker.cleanup.enabled=true Work?

Absolutely—this is the go-to config for automatic cleanup of old job data on Workers. But you'll want to pair it with related settings to make it effective:
Add these to spark-defaults.conf:

spark.worker.cleanup.enabled true
spark.worker.cleanup.interval 3600  # Run cleanup every hour (default)
spark.worker.cleanup.appDataTtl 86400  # Keep job data for 1 day (default)—older data gets deleted

Important note: This only cleans up data from completed/failed jobs, not currently running ones. So it won't interfere with active workloads, just clear out old clutter taking up space.

内容的提问来源于stack exchange,提问作者LennyMarshal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 22:42:59