Cloudera VM中Hadoop作业仅显示Reducer日志的问题求助
Hey there! I totally get the frustration of not seeing your Mapper and Partitioner logs in your Cloudera VM—been there, done that. Let's walk through the most common fixes to get those logs showing up:
First, make sure your log4j.properties (or log4j.xml) is properly set up to capture logs from your Mapper and Partitioner classes, and that it's being picked up by Hadoop. Here's a sample config snippet tailored to your package:
# Target your specific Mapper and Partitioner classes log4j.logger.org.apress.prohadoop.c6.YourMapperClass=INFO, console log4j.logger.org.apress.prohadoop.c6.YourPartitionerClass=INFO, console # Basic appender setup (adjust as needed) log4j.appender.console=org.apache.log4j.ConsoleAppender log4j.appender.console.layout=org.apache.log4j.PatternLayout log4j.appender.console.layout.ConversionPattern=%d{ISO8601} %-5p %c{1} - %m%n
Don't forget to:
- Package this config file inside your job JAR, or
- Specify its path explicitly when submitting the job with
-Dlog4j.configuration=file:/path/to/your/log4j.properties
Cloudera relies on YARN's log aggregation to collect task logs (including Mapper/Partitioner) and make them accessible in the UI. Here's how to check:
- Log into Cloudera Manager, navigate to the YARN service
- Go to the Configuration tab, search for these settings:
yarn.log-aggregation-enable: Must be set totrueyarn.log-aggregation.retain-seconds: Set to a value that gives you enough time to view logs (e.g., 86400 for 24 hours)
- If you change any settings, restart the YARN service to apply them
Mapper and Partitioner logs aren't always at the top-level "Logs" section of your job. You need to dig into the task attempts:
- Open your job's detail page in the Hadoop UI
- Switch to the Attempts tab
- Click on the specific Mapper attempt (each Mapper task has one or more attempts)
- The logs for that attempt will include both Mapper and Partitioner output, since Partitioner runs as part of the Mapper task lifecycle
Cloudera Manager sometimes enforces default log levels that might suppress your INFO-level logs. Check these settings:
- In Cloudera Manager, go to the MapReduce (or YARN) service's Configuration tab
- Search for:
mapreduce.map.log.level: Set toINFO(or your desired level, notERROR/WARN)mapreduce.reduce.log.level: Ensure this is also set appropriately
- Restart the relevant services if you make changes
Make sure you're actually using the logger correctly in your Mapper and Partitioner classes. Here's a quick example for your Mapper:
import org.apache.log4j.Logger; public class YourMapper extends Mapper<LongWritable, Text, Text, IntWritable> { private static final Logger LOG = Logger.getLogger(YourMapper.class); @Override protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException { LOG.info("Processing input line: " + value.toString()); // Your business logic here } }
Do the same for your Partitioner class—if you don't call LOG.info() (or other log methods) in the code, you won't see any output, no matter how good your config is.
If all else fails, force the logging settings when submitting your job with these command-line arguments:
hadoop jar your-job-file.jar org.apress.prohadoop.c6.YourMainJobClass \ -Dlog4j.configuration=classpath:log4j.properties \ -Dmapreduce.map.log.level=INFO \ -Dmapreduce.reduce.log.level=INFO
If your log4j config is outside the JAR, replace classpath:log4j.properties with file:/absolute/path/to/log4j.properties instead.
内容的提问来源于stack exchange,提问作者vibhas

