You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu下运行Hadoop3.1.0 MapReduce程序遇INFO提示求助

Fixing Excessive INFO Logs When Running Hadoop 3.1.0 MapReduce Jobs

Hey there! Let's sort out that overwhelming flood of INFO logs you're seeing when running your MapReduce job. I've run into this exact headache with Hadoop 3.x before, so here's a step-by-step breakdown of what to adjust:

1. Temporarily Reduce Log Output (Quick Test)

If you don't want to mess with config files right away, you can override the log level directly in your run command. This lets you see if your job actually runs without drowning in debug messages:

master@kevin-VirtualBox:~/MapReduceTutorial$ $HADOOP_HOME/bin/hadoop jar ProductSalePerCountry.jar /inputMapReduce /mapreduce_output_sales -Dlog4j.rootLogger=WARN,console

This switches the root logger to only show WARN and ERROR messages instead of every single INFO line.

2. Permanently Adjust Hadoop Log Configuration

For a long-term fix, edit the log4j config file in your Hadoop directory:

  • Navigate to $HADOOP_HOME/etc/hadoop/
  • Open log4j.properties (if you're on a newer Hadoop build, it might be log4j2.xml—check which one exists)

Make these changes to trim down the noise:

# Change the root logger level from INFO to WARN
log4j.rootLogger=WARN,console

# If you still want some MapReduce-specific info but less overall, tweak these instead:
log4j.logger.org.apache.hadoop.mapreduce=WARN
log4j.logger.org.apache.hadoop.yarn=WARN
log4j.logger.org.apache.hadoop.hdfs=WARN

Save the file, and restart your Hadoop services (stop-all.sh then start-all.sh) for changes to take effect.

3. Check Core Cluster Configs (Why the Job Might Be Stuck)

If the INFO logs are looping because your job isn't progressing, double-check these key config files:

core-site.xml

Ensure your HDFS address is correctly set for your cluster:

<property>
  <name>fs.defaultFS</name>
  <value>hdfs://master:9000</value>
</property>

hdfs-site.xml

Since you're running on a single-node VM, set replication to 1 (default is 3, which will cause copy failures):

<property>
  <name>dfs.replication</name>
  <value>1</value>
</property>

yarn-site.xml

Verify YARN has enough resources to run your job (adjust values based on your VM's available memory):

<property>
  <name>yarn.nodemanager.resource.memory-mb</name>
  <value>3072</value> <!-- Set to ~75% of your VM's total RAM -->
</property>
<property>
  <name>yarn.scheduler.maximum-allocation-mb</name>
  <value>3072</value> <!-- Match the above value -->
</property>

mapred-site.xml

Ensure MapReduce is configured to use YARN:

<property>
  <name>mapreduce.framework.name</name>
  <value>yarn</value>
</property>
<property>
  <name>mapreduce.map.memory.mb</name>
  <value>1024</value>
</property>
<property>
  <name>mapreduce.reduce.memory.mb</name>
  <value>2048</value>
</property>

4. Verify Cluster Services Are Running

Check your jps output to confirm all required daemons are up:

  • NameNode, DataNode (for HDFS)
  • ResourceManager, NodeManager (for YARN)
    If any are missing, restart the corresponding service with $HADOOP_HOME/sbin/hadoop-daemon.sh start <service-name>.

5. Use YARN WebUI to Diagnose Stuck Jobs

Open the YARN ResourceManager UI (default: http://master:8088) to check your job's status. If it's stuck in "Accepted" or "Running" without progress, it's likely a resource allocation issue (your VM doesn't have enough RAM/CPU) or input file problems (e.g., missing files in /inputMapReduce).

Give these steps a try—most of the time, the INFO log flood is either a log level setting or a stuck job due to misconfigured cluster resources.

内容的提问来源于stack exchange,提问作者Kevin Su

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:06:51