Python实现单词音节Top10统计报TypeError模块不可调用错误
问题背景
正在开发基于mrjob的MapReduce任务,目标是读取文本文件、统计每个单词的音节数量,最终返回音节数最多的前10个单词。完成大部分逻辑编写后运行抛出如下错误:
File "top_10_syllable_count.py", line 84, in get_syllable_count_pair return (syllables(word), word, ) TypeError: 'module' object is not callable
原始问题代码
import re from sys import stderr from mrjob.job import MRJob from mrjob.step import MRStep WORD_RE = re.compile(r"[\w']+") import syllables class MRMostUsedWordSyllables(MRJob): def steps(self): return [ MRStep(mapper=self.word_splitter_mapper, reducer=self.sorting_word_syllables), MRStep(mapper=self.get_syllable_count_pair), MRStep(reducer=self.get_top_10_reducer) ] def word_splitter_mapper(self, _, line): #for word in line.split(): for word in WORD_RE.findall(line): yield(word.lower(), None) def sorting_word_syllables(self, word, count): count = 0 vowels = 'aeiouy' word = word.lower().strip() if word in vowels: count +=1 for index in range(1,len(word)): if word[index] in vowels and word[index-1] not in vowels: count +=1 if word.endswith('e'): count -= 1 if word.endswith('le'): count+=1 if count == 0: count +=1 yield None, (int(count), word) def get_syllable_count_pair(self, _, word): return (syllables(word), word, ) def get_top_10_reducer(self, count, word): assert count == None # added for a guard with_counts = [get_syllable_count_pair(w) for w in word] # Sort the words by the syllable count sorted_counts = sorted(syllables_counts, reverse=True, key=lambda x: x[0]) # Slice off the first ten for t in sorted_counts[:10]: yield t if __name__ == '__main__': import time start = time.time() MRMostUsedWordSyllables.run() end = time.time() print(end - start)
错误原因
- 核心触发报错的原因:
import syllables导入的是整个syllables第三方模块,模块对象本身不能作为函数直接调用。syllables库提供的音节计数方法是模块下的estimate()函数,直接写syllables(word)属于调用模块而非调用功能函数,才会抛出module object is not callable错误。 - 代码中还存在其他会导致运行失败的逻辑问题:
- 类的实例方法
get_syllable_count_pair在reducer中被直接裸调用,没有加self.前缀,会触发名称错误 - 方法参数接收逻辑和上一个步骤的输出结构不匹配,
sorting_word_syllables输出的value是(音节数, 单词)的元组,传入下一个mapper时不能直接拆成单独的word参数 - 排序时使用了未定义的变量
syllables_counts,实际前面赋值的列表变量名为with_counts - 音节计数逻辑重复:既在
sorting_word_syllables中手动实现了音节计数规则,又额外调用syllables库做计数,逻辑冗余
- 类的实例方法
修复方法
按以下步骤调整代码即可正常运行:
- 把
syllables(word)的错误调用改成syllables.estimate(word),如果不需要用第三方库计数,也可以直接保留自己写的计数逻辑,删掉重复的syllables调用 - 调整MR各步骤的参数接收逻辑,对齐MapReduce的key-value数据流
- 修正类方法调用方式、变量名拼写错误
- 去掉重复的计数逻辑,简化步骤流程
修复后可运行代码
import re import time from mrjob.job import MRJob from mrjob.step import MRStep WORD_RE = re.compile(r"[\w']+") import syllables class MRMostUsedWordSyllables(MRJob): def steps(self): return [ MRStep(mapper=self.word_splitter_mapper, reducer=self.count_syllables_reducer), MRStep(reducer=self.get_top_10_reducer) ] def word_splitter_mapper(self, _, line): for word in WORD_RE.findall(line): # 去重后统计,避免重复单词重复计算 yield word.lower(), 1 def count_syllables_reducer(self, word, _): # 调用syllables库的正确方法统计音节数 syllable_count = syllables.estimate(word) # 所有结果输出到同一个key下,方便后续全局排序 yield None, (syllable_count, word) def get_top_10_reducer(self, _, count_word_pairs): # 按音节数倒序排序,取前10 sorted_pairs = sorted(count_word_pairs, reverse=True, key=lambda x: x[0]) for pair in sorted_pairs[:10]: yield pair if __name__ == '__main__': start = time.time() MRMostUsedWordSyllables.run() end = time.time() print(f"任务运行耗时: {end - start}s")
注:如果不想依赖第三方syllables库,把
count_syllables_reducer里的syllables.estimate(word)替换成原来手动写的音节计数逻辑即可。
内容的提问来源于stack exchange,提问作者Tony M
相关产品推荐
相关产品推荐

