Hadoop集群内存耗尽重启后Spark服务start-all.sh启动失败求助
Hey there, let’s work through this cluster restart issue step by step. The forced shutdown from memory exhaustion almost certainly left behind residual state (zombie processes, lock files, corrupted temp data) that’s messing with your ./start-all.sh script. Here’s how to fix it:
1. Wipe Residual Processes & Temporary Files
First, we need to clear out any leftover processes and stale temporary data that might be holding locks or conflicting with new service starts:
- Kill all Hadoop/Spark-related processes (even hidden zombie ones):
ps aux | grep -E 'hadoop|spark|java' | grep -v grep | awk '{print $2}' | xargs sudo kill -9 - Delete temporary directories used by Hadoop and Spark (these get recreated automatically):
rm -rf /tmp/hadoop-* /tmp/spark-* /tmp/yarn-* - Check for leftover lock files in your Hadoop data/name node directories. Look for files named
in_use.lockunder paths defined inhdfs-site.xml(e.g.,dfs.name.dir,dfs.data.dir) and delete them if present.
2. Fix File Permissions
Forced shutdowns and user context changes can mess up directory permissions, especially for log paths like /dbhome:
- Verify permissions on your log directory:
ls -ld /dbhome/ - If the directory isn’t owned by your Hadoop/Spark user, fix it:
(Replacesudo chown -R your-hadoop-user:your-hadoop-group /dbhome/your-hadoop-userandyour-hadoop-groupwith your actual cluster user/group, e.g.,hadoop:hadoop.)
3. Start Services One-by-One (Avoid start-all.sh Temporarily)
The start-all.sh script hides individual service errors. Let’s isolate which component is failing by starting services manually:
- Start Hadoop HDFS first:
Check logs for errors (look inhdfs --daemon start namenode hdfs --daemon start datanode$HADOOP_HOME/logsor your custom log path) before moving on. - Start YARN:
yarn --daemon start resourcemanager yarn --daemon start nodemanager - Start Spark components separately (since you confirmed Master works alone):
This will tell you if the Worker node is the one failing when using./start-master.sh # Wait 30 seconds, then start workers with explicit Master URL ./start-worker.sh spark://your-master-ip:7077start-all.sh.
4. Adjust Memory Configurations (Prevent Repeat Memory Exhaustion)
Since your cluster crashed from memory exhaustion, tweak your Spark/Hadoop configs to avoid overcommitting:
- In
spark-env.sh, set a reasonable worker memory limit (match to your machine’s available RAM):export SPARK_WORKER_MEMORY=8g # Adjust based on your node's total RAM export SPARK_DRIVER_MEMORY=4g - In Hadoop’s
yarn-site.xml, lower the container memory limits:<property> <name>yarn.nodemanager.resource.memory-mb</name> <value>16384</value> # Set to ~80% of your node's total RAM </property>
5. Dig Into Detailed Logs
The truncated log path you mentioned (/dbhome...) holds the key to exact errors. Pull the full error stack trace from the relevant log file:
tail -n 100 /dbhome/spark/logs/spark-*-master-*.out
Look for keywords like Address already in use, Permission denied, or Corrupted metadata—these will point you to the root cause if the above steps don’t resolve it.
内容的提问来源于stack exchange,提问作者dhanush-ai1990

