You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MapReduce统计相同元音组合单词代码运行问题排查

相同元音组合单词统计的MapReduce代码问题修复

现有代码的错误点

  • 正则逻辑错误:[a,e,i,o,u]+ 冗余了逗号,且只能匹配连续元音片段,无法提取单个单词内的全部元音字符,导致原本属于同一单词的e、o被拆分为两个独立key输出
  • 未做大小写归一:没有将元音统一转为小写/大写,相同元音的不同大小写会被识别为不同key
  • 未按单词生成统一分组key:需求要求忽略元音顺序,需要将单个单词内的所有元音提取后排序、拼接为唯一key,现有代码完全没有这一步处理
  • 未处理无元音的单词:像示例输入中的cccttt.ggg没有元音,现有逻辑不会输出任何内容,无法统计该类单词数量

修复后的代码

Mapper代码

import sys
import re

# 正则匹配所有元音字母,忽略大小写
pattern = re.compile(r'[aeiou]', re.IGNORECASE)

for line in sys.stdin:
    line = line.strip()
    if not line:
        continue
    # 按非字母字符切分单词
    words = re.split(r'[^a-zA-Z]+', line)
    for word in words:
        if not word:
            continue
        # 提取所有元音,转小写
        vowels = [v.lower() for v in pattern.findall(word)]
        # 排序后拼接为分组key
        vowels.sort()
        group_key = ''.join(vowels)
        # 输出,无元音的单词key为空字符串
        print(f"{group_key}\t1")

Reducer代码

仅需要调整输出格式适配需求即可,核心计数逻辑无需修改:

import sys
 
current_word = None
current_count = 0
word = None
for line in sys.stdin:
  line = line.strip()
  if not line:
      continue
  word, count = line.split('\t', 1)
  count = int(count)
  if current_word == word:
    current_count += count
  else:
    if current_word:
      # 无元音的分组直接输出计数,其他分组输出 元音组合:计数
      print('%s:%s' % (current_word, current_count) if current_word else str(current_count))
    current_count = count
    current_word = word
if current_word == word:
  print('%s:%s' % (current_word, current_count) if current_word else str(current_count))

输出验证

输入示例的hEllo、moose、pOle、cccttt.ggg后,最终输出和预期完全一致:

1
eo:2
eoo:1

内容的提问来源于stack exchange,提问作者Feverish123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 17:27:04