You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python MapReduce概率计算异常:Google DataProc集群运行问题排查

Fixing MapReduce Probability Calculation for CSV Data on Google DataProc

Let's break down what's going wrong with your MapReduce code and fix it step by step:

1. Core Issues in Your Current Code

Mapper Problem

Your CSV line is probability_mass 10 10 60 10 30, but in mapper.py, you're using for probability_mass in line — this iterates over each individual character in the line (like '1', '0', ' ', '6', etc.), not the full numeric values you care about. That's exactly why your reducer is counting single digits instead of complete numbers.

Reducer Minor Tweak

Your reducer's logic is almost there, but we'll clean it up to align with the fixed mapper output and match your expected format better.

2. Corrected Mapper Code (mapper.py)

#!/usr/bin/env python3
"""mapper.py"""
import sys

# Process each input line
for line in sys.stdin:
    line = line.strip()
    # Skip empty lines to avoid errors
    if not line:
        continue
    # Split the line into parts: first element is the header, rest are values
    parts = line.split()
    # Skip the header "probability_mass" and loop through numeric values
    for value in parts[1:]:
        # Emit a fixed key + value pair to group all numbers in the reducer
        print(f"count\t{value}")
  • We split the line using split() (which handles any whitespace by default)
  • We skip the first element (parts[1:]) since it's just the header text
  • We use a fixed key (count) to ensure all numeric values are grouped together in the reducer

3. Corrected Reducer Code (reducer.py)

#!/usr/bin/env python3
"""reducer.py"""
import sys
from collections import defaultdict

counts = defaultdict(int)

# Read all input from the mapper
for line in sys.stdin:
    line = line.strip()
    if not line:
        continue
    # Parse the key-value pair (we can ignore the key since it's fixed)
    _, v = line.split('\t', 1)
    counts[v] += 1

# Calculate total number of values processed
total = sum(counts.values())
# Compute probabilities, rounded to 1 decimal place to match your expected output
probabilities = {k: round(v / total, 1) for k, v in counts.items()}
# Print the result in the desired format
print(probabilities)
  • We ignore the key from the mapper since we only need to count all numeric values
  • We round probabilities to one decimal place to match your expected output (0.6 instead of 0.600000...)
  • The output will now exactly match the format you're looking for

4. Expected Output After Fixes

Once you run the corrected code with your existing Hadoop Streaming command, the output in /tmp/output should be:

{'10': 0.6, '60': 0.2, '30': 0.2}

Quick Additional Notes

  • This code will handle multiple lines in your CSV (if you expand beyond one row) — it skips the header on each line and aggregates all numeric values across the dataset
  • The empty line checks prevent errors from stray blank lines in your input CSV

内容的提问来源于stack exchange,提问作者Humair Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:54:22