Python MapReduce概率计算异常:Google DataProc集群运行问题排查
Let's break down what's going wrong with your MapReduce code and fix it step by step:
1. Core Issues in Your Current Code
Mapper Problem
Your CSV line is probability_mass 10 10 60 10 30, but in mapper.py, you're using for probability_mass in line — this iterates over each individual character in the line (like '1', '0', ' ', '6', etc.), not the full numeric values you care about. That's exactly why your reducer is counting single digits instead of complete numbers.
Reducer Minor Tweak
Your reducer's logic is almost there, but we'll clean it up to align with the fixed mapper output and match your expected format better.
2. Corrected Mapper Code (mapper.py)
#!/usr/bin/env python3 """mapper.py""" import sys # Process each input line for line in sys.stdin: line = line.strip() # Skip empty lines to avoid errors if not line: continue # Split the line into parts: first element is the header, rest are values parts = line.split() # Skip the header "probability_mass" and loop through numeric values for value in parts[1:]: # Emit a fixed key + value pair to group all numbers in the reducer print(f"count\t{value}")
- We split the line using
split()(which handles any whitespace by default) - We skip the first element (
parts[1:]) since it's just the header text - We use a fixed key (
count) to ensure all numeric values are grouped together in the reducer
3. Corrected Reducer Code (reducer.py)
#!/usr/bin/env python3 """reducer.py""" import sys from collections import defaultdict counts = defaultdict(int) # Read all input from the mapper for line in sys.stdin: line = line.strip() if not line: continue # Parse the key-value pair (we can ignore the key since it's fixed) _, v = line.split('\t', 1) counts[v] += 1 # Calculate total number of values processed total = sum(counts.values()) # Compute probabilities, rounded to 1 decimal place to match your expected output probabilities = {k: round(v / total, 1) for k, v in counts.items()} # Print the result in the desired format print(probabilities)
- We ignore the key from the mapper since we only need to count all numeric values
- We round probabilities to one decimal place to match your expected output (
0.6instead of0.600000...) - The output will now exactly match the format you're looking for
4. Expected Output After Fixes
Once you run the corrected code with your existing Hadoop Streaming command, the output in /tmp/output should be:
{'10': 0.6, '60': 0.2, '30': 0.2}
Quick Additional Notes
- This code will handle multiple lines in your CSV (if you expand beyond one row) — it skips the header on each line and aggregates all numeric values across the dataset
- The empty line checks prevent errors from stray blank lines in your input CSV
内容的提问来源于stack exchange,提问作者Humair Ali

