You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将元组列表归约为单个输出?MapReduce归约器开发求助

问题描述

已完成mapper和groupby脚本,终端运行后得到以下格式的输入:

a [('1', '0', '0'), ('1', '1', '1')]
b [('1', '0', '0')]
e [('1', '0', '0'), ('1', '0', '1')]
f [('1', '0', '0'), ('1', '0', '0')]
i [('1', '0', '0'), ('1', '0', '0'), ('1', '1', '0')]
l [('1', '0', '0'), ('1', '0', '0')]
s [('1', '0', '0')]
t [('1', '0', '0'), ('1', '0', '0')]
u [('1', '0', '0'), ('1', '0', '0')]

需要编写reducer脚本,接收每行输入后,将每个元组对应索引的元素求和,输出指定格式:

a 2 1 1
b 1 0 0
e 2 0 1
f 2 0 0
i 3 0 0
l 2 0 0
t 1 0 0
u 2 0 0

修改示例代码后运行出现错误:

ValueError: invalid literal for int() with base 10: "[('1',"

现有代码如下:

import sys


#The function read_mapper_output is similar to the function in group-by-key.py
def read_mapper_output(file):
    for line in file:
        yield line.strip().split(' ')

#Read the input one by one
for vec in read_mapper_output(sys.stdin):
    #The first element after split is the word
    word = vec[0]
    #Count the number of ones
    count = sum(int(number) for number in vec[1:])
    #Print the word followed by number
    print("%s %d" % (word, count))
问题原因与解决方法

问题根源

用split(' ')按空格分割输入行时,会把元组列表的结构拆成零散的字符串片段(比如"[('1',"、"'0',"这类),这些字符串无法直接转成整数,导致报错。同时原代码仅做了总和计算,没有按索引分别求和,不符合输出要求。

修复步骤

  1. 解析元组列表字符串:输入行后半部分是合法的Python元组列表字符串,用ast.literal_eval()直接解析成真实列表对象,避免手动分割的麻烦。
  2. 按索引分别求和:初始化长度为3的列表(对应每个元组的3个元素),遍历每个元组,将对应索引的元素转成整数后累加。
  3. 格式化输出:将单词和三个求和结果按空格拼接输出。

修复后的代码

import sys
import ast

def read_mapper_output(file):
    for line in file:
        # 仅按第一个空格分割,分离单词和元组列表字符串
        word, tuples_str = line.strip().split(' ', 1)
        # 解析字符串为可操作的元组列表
        tuples_list = ast.literal_eval(tuples_str)
        yield word, tuples_list

for word, tuples_list in read_mapper_output(sys.stdin):
    # 初始化三个索引的求和值为0
    sums = [0, 0, 0]
    for tpl in tuples_list:
        sums[0] += int(tpl[0])
        sums[1] += int(tpl[1])
        sums[2] += int(tpl[2])
    # 按要求格式输出
    print(f"{word} {' '.join(map(str, sums))}")

代码说明

  • split(' ', 1)只执行一次分割,确保单词和完整的元组列表字符串被正确分离,不会破坏元组结构。
  • ast.literal_eval()安全地将字符串形式的元组列表转换为Python可操作的对象,比eval()更安全,避免代码注入风险。
  • 遍历元组时分别对三个索引位置的元素累加,最终得到对应位置的总和,匹配指定输出格式。

内容的提问来源于stack exchange,提问作者Zaku

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 18:42:43