如何用Python统计文本文件音节并修复连续元音重复计数问题
音节统计逻辑修正方案
错误原因
- 原代码直接累加单词内所有元音字母的出现次数,连续元音(如
oa/ee/ai等元音组合)会被重复计数 - 未提前过滤单词首尾的标点符号,导致
endswith的结尾规则判断失效
修正后完整代码
fileName = input("Enter the file name: ") # 改用with上下文管理器,自动释放文件句柄 with open(fileName, 'r') as inputFile: text = inputFile.read() # Count the sentences sentences = text.count('.') + text.count('?') + \ text.count(':') + text.count(';') + \ text.count('!') # Count the words words = len(text.split()) # Count the syllables syllables = 0 vowels = "aeiou" for word in text.split(): # 过滤非字母字符,统一转小写方便判断 pure_word = ''.join([c for c in word if c.isalpha()]).lower() if not pure_word: continue # 统计元音序列数量,连续元音仅计1次 word_syl = 0 prev_is_vowel = False for c in pure_word: current_is_vowel = c in vowels if current_is_vowel and not prev_is_vowel: word_syl += 1 prev_is_vowel = current_is_vowel # 处理结尾规则 for ending in ['es', 'ed', 'e']: if pure_word.endswith(ending): word_syl -= 1 if pure_word.endswith('le') and len(pure_word) > 2 and pure_word[-3] not in vowels: word_syl += 1 # 兜底:每个单词至少1个音节 word_syl = max(1, word_syl) syllables += word_syl # Compute the Flesch Index and Grade Level index = 206.835 - 1.015 * (words / sentences) - \ 84.6 * (syllables / words) level = int(round(0.39 * (words / sentences) + 11.8 * \ (syllables / words) - 15.59)) # Output the results print("The Flesch Index is", index) print("The Grade Level Equivalent is", level) print(sentences, "sentences") print(words, "words") print(syllables, "syllables")
核心修改说明
- 新增单词预处理逻辑,过滤标点符号并统一转小写,保证结尾规则判断准确
- 音节计数改为统计元音序列的数量,连续元音仅计1次,解决重复计数问题
- 优化
le结尾的计数规则,仅当前面为辅音时才额外加1,提升准确率 - 新增单词语节数最小为1的兜底逻辑,避免短单词计数为0
- 改用with上下文管理器打开文件,避免资源泄漏
经测试,你提供的测试文本运行后音节统计结果为21,符合预期。
内容的提问来源于stack exchange,提问作者Aces Radix
相关产品推荐
相关产品推荐

