You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python mrjob提取10个最长单词结果重复,如何去重获取唯一值?

解决mrjob查找最长单词时的重复问题

问题原因

现有代码会将文本中每一次出现的单词都生成(单词长度, 小写单词)元组提交给reducer,同一个单词多次出现就会生成多条相同元组,排序后输出自然会有重复条目。

解决方法

在reducer排序前先对收到的元组去重即可,修改后的完整代码如下:

%%file most_chars.py  
from mrjob.job import MRJob
from mrjob.step import MRStep
import re

WORD_RE = re.compile(r"[\w']+") # 匹配任意字母数字或撇号,用于拆分每行文本


class MostChars(MRJob):
    def steps(self):
        return [
            MRStep(mapper=self.mapper_get_words,
                  reducer=self.reducer_find_longest_words)
        ]

    def mapper_get_words(self, _, line):
        for word in WORD_RE.findall(line):   
            yield None, (len(word), word.lower().strip())

    # 忽略键,其值仅为None
    def reducer_find_longest_words(self, _, word_count_pairs):
        # 先对(长度,单词)元组去重,同一个单词仅保留一条记录
        unique_pairs = list({pair for pair in word_count_pairs})
        # 对去重后的记录按单词长度倒序排序
        sorted_pair = sorted(unique_pairs, reverse=True)
        # 取前10个最长的唯一单词输出
        for pair in sorted_pair[0:10]:
            yield pair
              
if __name__ == '__main__':
    MostChars.run()

可选优化

如果处理的文本体量很大,可以额外加combiner阶段在map端先做局部去重,减少shuffle阶段的数据传输量:

def steps(self):
    return [
        MRStep(mapper=self.mapper_get_words,
              combiner=self.combiner_deduplicate,
              reducer=self.reducer_find_longest_words)
    ]

def combiner_deduplicate(self, _, pairs):
    for pair in set(pairs):
        yield None, pair

内容的提问来源于stack exchange,提问作者Nekojell

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 01:39:05