Hadoop DataNode启动失败(返回值1)求助:主从VM部署异常
Hey there, let's figure out why your DataNode won't start on the slave node with that return value 1 error. This is a generic failure code—so we need to dig into the specifics to fix it. Here are the most common troubleshooting steps I recommend:
First, check the DataNode logs for detailed errors
The return value 1 doesn't tell us much, but the logs will have the real issue. By default, Hadoop logs live in$HADOOP_HOME/logs/—look for a file named something likehadoop-<your-username>-datanode-<slave-hostname>.log. Run this to see the latest entries:tail -n 50 $HADOOP_HOME/logs/hadoop-$(whoami)-datanode-$(hostname).logYou'll likely find clues here—like permission issues, mismatched cluster IDs, or port conflicts.
Verify configuration files match between master and slave
All core configs on the slave must be identical to the master. Double-check these files:$HADOOP_HOME/etc/hadoop/core-site.xml: Make surefs.defaultFSpoints to your master's NameNode address (e.g.,hdfs://master:9000)$HADOOP_HOME/etc/hadoop/hdfs-site.xml: Confirmdfs.datanode.data.dirpoints to an existing path, anddfs.namenode.rpc-addressis set correctly to the master$HADOOP_HOME/etc/hadoop/workers: Ensure the slave's hostname/IP is listed here, and the master can SSH to the slave without a password
Check directory permissions and existence
The DataNode needs read/write access to its data directory and log directory:- First, confirm the data directory from
hdfs-site.xmlexists:DATA_DIR=$(grep dfs.datanode.data.dir $HADOOP_HOME/etc/hadoop/hdfs-site.xml | awk -F '>' '{print $2}' | awk -F '<' '{print $1}') ls -ld $DATA_DIR - If it's missing, create it. Then make sure it's owned by the Hadoop user (e.g.,
hadoop):sudo chown -R hadoop:hadoop $DATA_DIR sudo chmod -R 755 $DATA_DIR
- First, confirm the data directory from
Fix mismatched Cluster IDs
If you formatted the NameNode on the master after setting up the slave, the DataNode might have an old cluster ID that doesn't match. Here's how to check and fix:- On the master, get the NameNode's cluster ID:
cat $HADOOP_HOME/dfs/name/current/VERSION | grep clusterID - On the slave, check the DataNode's cluster ID:
cat $HADOOP_HOME/dfs/data/current/VERSION | grep clusterID - If they don't match, delete the slave's data directory (warning: this erases all HDFS data on the slave!) and restart the DataNode:
The DataNode will automatically sync the correct cluster ID from the master on startup.rm -rf $HADOOP_HOME/dfs/data hdfs datanode
- On the master, get the NameNode's cluster ID:
Check for port conflicts
DataNodes use default ports like 50010 and 50020. Make sure these aren't taken by other processes:netstat -tulpn | grep 50010 netstat -tulpn | grep 50020If a port is occupied, either stop the conflicting process or update the port numbers in
hdfs-site.xml.
Once you've worked through these steps, try starting the DataNode again with hdfs datanode or start-dfs.sh from the master. Let me know if the logs show a specific error you need help interpreting!
内容的提问来源于stack exchange,提问作者elle believe

