Python文本生成器多字符统计时出现KeyError错误求助
3字符链统计文本生成时的KeyError问题
问题描述
写代码统计文本字符序列并生成新文本,单字符统计功能正常,但换成3字符链统计时,出现KeyError错误,提示找不到键' -–'。
相关代码
import zipfile from random import randint from pprint import pprint # file_zip = 'voyna-i-mir.txt.zip' # zip = zipfile.ZipFile(file_zip, 'r') # for file in zip.namelist(): # zip.extract(file) origin = 'voyna-i-mir.txt' statistic = {} chain = ' ' with open(origin, 'r', encoding='cp1251') as file: for sting in file: #print(sting) for symbol in sting: if chain in statistic: if symbol in statistic[chain]: statistic[chain][symbol] += 1 else: statistic[chain][symbol] = 1 else: statistic[chain] = {symbol: 1} chain = chain[1:] + symbol dictionary = {} stat_generator = {} for chain, symbol_stat in statistic.items(): dictionary[chain] = 0 stat_generator[chain] = [] for symbol, count in symbol_stat.items(): dictionary[chain] += count stat_generator[chain].append([count, symbol]) stat_generator[chain].sort(reverse=True) gen = 1000 was_print = 0 chain = ' ' while was_print < gen: symbol_stat = stat_generator[chain] total = dictionary[chain] random = randint(1, total) position = 0 for count, symbol in symbol_stat: position += count if random <= position: break print(symbol, end='') was_print += 1 chain = chain[1:] + symbol
错误信息
Traceback (most recent call last): File "C:\Users\roman\Desktop\skillbox\[Skillbox] Профессия Python- разработчик\9. Работа с файлами и форматированный вывод-20210102T191733Z-001\9. Работа с файлами и форматированный вывод\lesson_009\python_snippets\test.py", line 46, in <module> symbol_stat = stat_generator[chain] ~~~~~~~~~~~~~~^^^^^^^ KeyError: ' -–'
问题根源
- 生成逻辑完全颠倒:代码里
print和更新chain的操作写在了break的else分支,导致找到目标字符前就错误更新了链,生成原文本中从未出现过的3字符组合,自然在统计字典里找不到对应键。 - 无容错处理:即使逻辑正确,生成过程中也可能出现原文本未覆盖的链组合,此时直接去字典取值就会触发KeyError。
修复方案
1. 修正生成循环的核心逻辑
把字符输出和链更新的代码移到找到目标字符之后,确保每次更新的链都是基于合法统计结果生成的:
2. 增加链不存在的容错处理
当生成的链不在统计字典中时,重置为初始链,避免报错。
修正后的生成循环代码:
gen = 1000 was_print = 0 chain = ' ' while was_print < gen: # 处理不存在的链,重置为初始状态 if chain not in stat_generator: chain = ' ' continue symbol_stat = stat_generator[chain] total = dictionary[chain] random = randint(1, total) position = 0 target_symbol = None # 找到随机命中的字符 for count, symbol in symbol_stat: position += count if random <= position: target_symbol = symbol break # 输出并更新链 if target_symbol: print(target_symbol, end='') was_print += 1 chain = chain[1:] + target_symbol
额外优化建议
- 修正变量名拼写:
for sting in file:改为for string in file:,提升代码可读性。 - 用
collections.defaultdict简化统计逻辑,省去多层if判断:
from collections import defaultdict statistic = defaultdict(lambda: defaultdict(int)) chain = ' ' with open(origin, 'r', encoding='cp1251') as file: for string in file: for symbol in string: statistic[chain][symbol] += 1 chain = chain[1:] + symbol
内容的提问来源于stack exchange,提问作者О. К.
相关产品推荐
相关产品推荐

