MapReduce统计相同元音组合单词代码运行问题排查
相同元音组合单词统计的MapReduce代码问题修复
现有代码的错误点
- 正则逻辑错误:
[a,e,i,o,u]+冗余了逗号,且只能匹配连续元音片段,无法提取单个单词内的全部元音字符,导致原本属于同一单词的e、o被拆分为两个独立key输出 - 未做大小写归一:没有将元音统一转为小写/大写,相同元音的不同大小写会被识别为不同key
- 未按单词生成统一分组key:需求要求忽略元音顺序,需要将单个单词内的所有元音提取后排序、拼接为唯一key,现有代码完全没有这一步处理
- 未处理无元音的单词:像示例输入中的
cccttt.ggg没有元音,现有逻辑不会输出任何内容,无法统计该类单词数量
修复后的代码
Mapper代码
import sys import re # 正则匹配所有元音字母,忽略大小写 pattern = re.compile(r'[aeiou]', re.IGNORECASE) for line in sys.stdin: line = line.strip() if not line: continue # 按非字母字符切分单词 words = re.split(r'[^a-zA-Z]+', line) for word in words: if not word: continue # 提取所有元音,转小写 vowels = [v.lower() for v in pattern.findall(word)] # 排序后拼接为分组key vowels.sort() group_key = ''.join(vowels) # 输出,无元音的单词key为空字符串 print(f"{group_key}\t1")
Reducer代码
仅需要调整输出格式适配需求即可,核心计数逻辑无需修改:
import sys current_word = None current_count = 0 word = None for line in sys.stdin: line = line.strip() if not line: continue word, count = line.split('\t', 1) count = int(count) if current_word == word: current_count += count else: if current_word: # 无元音的分组直接输出计数,其他分组输出 元音组合:计数 print('%s:%s' % (current_word, current_count) if current_word else str(current_count)) current_count = count current_word = word if current_word == word: print('%s:%s' % (current_word, current_count) if current_word else str(current_count))
输出验证
输入示例的hEllo、moose、pOle、cccttt.ggg后,最终输出和预期完全一致:
1 eo:2 eoo:1
内容的提问来源于stack exchange,提问作者Feverish123
相关产品推荐
相关产品推荐

