Ubuntu下运行Hadoop3.1.0 MapReduce程序遇INFO提示求助
Hey there! Let's sort out that overwhelming flood of INFO logs you're seeing when running your MapReduce job. I've run into this exact headache with Hadoop 3.x before, so here's a step-by-step breakdown of what to adjust:
1. Temporarily Reduce Log Output (Quick Test)
If you don't want to mess with config files right away, you can override the log level directly in your run command. This lets you see if your job actually runs without drowning in debug messages:
master@kevin-VirtualBox:~/MapReduceTutorial$ $HADOOP_HOME/bin/hadoop jar ProductSalePerCountry.jar /inputMapReduce /mapreduce_output_sales -Dlog4j.rootLogger=WARN,console
This switches the root logger to only show WARN and ERROR messages instead of every single INFO line.
2. Permanently Adjust Hadoop Log Configuration
For a long-term fix, edit the log4j config file in your Hadoop directory:
- Navigate to
$HADOOP_HOME/etc/hadoop/ - Open
log4j.properties(if you're on a newer Hadoop build, it might belog4j2.xml—check which one exists)
Make these changes to trim down the noise:
# Change the root logger level from INFO to WARN log4j.rootLogger=WARN,console # If you still want some MapReduce-specific info but less overall, tweak these instead: log4j.logger.org.apache.hadoop.mapreduce=WARN log4j.logger.org.apache.hadoop.yarn=WARN log4j.logger.org.apache.hadoop.hdfs=WARN
Save the file, and restart your Hadoop services (stop-all.sh then start-all.sh) for changes to take effect.
3. Check Core Cluster Configs (Why the Job Might Be Stuck)
If the INFO logs are looping because your job isn't progressing, double-check these key config files:
core-site.xml
Ensure your HDFS address is correctly set for your cluster:
<property> <name>fs.defaultFS</name> <value>hdfs://master:9000</value> </property>
hdfs-site.xml
Since you're running on a single-node VM, set replication to 1 (default is 3, which will cause copy failures):
<property> <name>dfs.replication</name> <value>1</value> </property>
yarn-site.xml
Verify YARN has enough resources to run your job (adjust values based on your VM's available memory):
<property> <name>yarn.nodemanager.resource.memory-mb</name> <value>3072</value> <!-- Set to ~75% of your VM's total RAM --> </property> <property> <name>yarn.scheduler.maximum-allocation-mb</name> <value>3072</value> <!-- Match the above value --> </property>
mapred-site.xml
Ensure MapReduce is configured to use YARN:
<property> <name>mapreduce.framework.name</name> <value>yarn</value> </property> <property> <name>mapreduce.map.memory.mb</name> <value>1024</value> </property> <property> <name>mapreduce.reduce.memory.mb</name> <value>2048</value> </property>
4. Verify Cluster Services Are Running
Check your jps output to confirm all required daemons are up:
- NameNode, DataNode (for HDFS)
- ResourceManager, NodeManager (for YARN)
If any are missing, restart the corresponding service with$HADOOP_HOME/sbin/hadoop-daemon.sh start <service-name>.
5. Use YARN WebUI to Diagnose Stuck Jobs
Open the YARN ResourceManager UI (default: http://master:8088) to check your job's status. If it's stuck in "Accepted" or "Running" without progress, it's likely a resource allocation issue (your VM doesn't have enough RAM/CPU) or input file problems (e.g., missing files in /inputMapReduce).
Give these steps a try—most of the time, the INFO log flood is either a log level setting or a stuck job due to misconfigured cluster resources.
内容的提问来源于stack exchange,提问作者Kevin Su

