You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hadoop集群内存耗尽重启后Spark服务start-all.sh启动失败求助

Troubleshooting Hadoop/Spark Cluster Restart Errors After Forced Shutdown

Hey there, let’s work through this cluster restart issue step by step. The forced shutdown from memory exhaustion almost certainly left behind residual state (zombie processes, lock files, corrupted temp data) that’s messing with your ./start-all.sh script. Here’s how to fix it:

1. Wipe Residual Processes & Temporary Files

First, we need to clear out any leftover processes and stale temporary data that might be holding locks or conflicting with new service starts:

  • Kill all Hadoop/Spark-related processes (even hidden zombie ones):
    ps aux | grep -E 'hadoop|spark|java' | grep -v grep | awk '{print $2}' | xargs sudo kill -9
    
  • Delete temporary directories used by Hadoop and Spark (these get recreated automatically):
    rm -rf /tmp/hadoop-* /tmp/spark-* /tmp/yarn-*
    
  • Check for leftover lock files in your Hadoop data/name node directories. Look for files named in_use.lock under paths defined in hdfs-site.xml (e.g., dfs.name.dir, dfs.data.dir) and delete them if present.

2. Fix File Permissions

Forced shutdowns and user context changes can mess up directory permissions, especially for log paths like /dbhome:

  • Verify permissions on your log directory:
    ls -ld /dbhome/
    
  • If the directory isn’t owned by your Hadoop/Spark user, fix it:
    sudo chown -R your-hadoop-user:your-hadoop-group /dbhome/
    
    (Replace your-hadoop-user and your-hadoop-group with your actual cluster user/group, e.g., hadoop:hadoop.)

3. Start Services One-by-One (Avoid start-all.sh Temporarily)

The start-all.sh script hides individual service errors. Let’s isolate which component is failing by starting services manually:

  1. Start Hadoop HDFS first:
    hdfs --daemon start namenode
    hdfs --daemon start datanode
    
    Check logs for errors (look in $HADOOP_HOME/logs or your custom log path) before moving on.
  2. Start YARN:
    yarn --daemon start resourcemanager
    yarn --daemon start nodemanager
    
  3. Start Spark components separately (since you confirmed Master works alone):
    ./start-master.sh
    # Wait 30 seconds, then start workers with explicit Master URL
    ./start-worker.sh spark://your-master-ip:7077
    
    This will tell you if the Worker node is the one failing when using start-all.sh.

4. Adjust Memory Configurations (Prevent Repeat Memory Exhaustion)

Since your cluster crashed from memory exhaustion, tweak your Spark/Hadoop configs to avoid overcommitting:

  • In spark-env.sh, set a reasonable worker memory limit (match to your machine’s available RAM):
    export SPARK_WORKER_MEMORY=8g  # Adjust based on your node's total RAM
    export SPARK_DRIVER_MEMORY=4g
    
  • In Hadoop’s yarn-site.xml, lower the container memory limits:
    <property>
        <name>yarn.nodemanager.resource.memory-mb</name>
        <value>16384</value>  # Set to ~80% of your node's total RAM
    </property>
    

5. Dig Into Detailed Logs

The truncated log path you mentioned (/dbhome...) holds the key to exact errors. Pull the full error stack trace from the relevant log file:

tail -n 100 /dbhome/spark/logs/spark-*-master-*.out

Look for keywords like Address already in use, Permission denied, or Corrupted metadata—these will point you to the root cause if the above steps don’t resolve it.


内容的提问来源于stack exchange,提问作者dhanush-ai1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 08:02:34