Python字母移位解码程序无法匹配文本字典行与字符串求助
解决字母移位解密中的字典匹配失效问题
你编写的Python程序用于遍历所有字母移位可能性,并与engmix.txt字典中的条目匹配,但遇到相同单词时if word == entry.strip()判定不生效,最终导致scores数组无有效数据。以下是问题分析和修复方案:
原代码
def autoDecode(userString): scores = [0 for i in range(25)] sentenceArray = decode(userString) score = 0 # loop over each sentence in the sentence array with open('engmix.txt') as dictionary: for sentence in sentenceArray: sentence = sentence.split(" ") sentencePos = 1 for word in sentence: word = word.lower() for entry in dictionary: if word == entry.strip(): score += 1 scores[sentencePos] = score score = 0 sentencePos = 0 dictionary.seek(0) print(scores) print("The sentence is likely to be", sentenceArray[sentencePos])
问题分析
- 字典遍历指针问题:每个单词循环时,字典会从当前指针位置开始读取,第一个单词读完字典后,指针停在文件末尾,后续单词无法再读取字典内容,导致匹配失效。虽然你在每个句子循环后用
seek(0)重置指针,但每个单词循环内部没有重置,只会在第一个单词时遍历字典,后续单词直接跳过匹配逻辑。 - 索引逻辑错误:
sentencePos初始设为1,每次句子循环后又重置为0,导致所有得分都被赋值到scores[1],最终scores数组大部分位置都是初始的0;最后输出时取sentenceArray[sentencePos]永远是第一个元素,逻辑完全错误。 - 匹配格式不一致:单词和字典条目可能存在大小写、多余空白的差异,仅对单词转小写无法完全覆盖格式问题。
修复方案
方案1:修正指针与索引逻辑(保留逐行匹配思路)
def autoDecode(userString): scores = [0] * 25 sentenceArray = decode(userString) with open('engmix.txt') as dictionary: # 遍历每个移位后的句子,用enumerate记录索引 for idx, sentence in enumerate(sentenceArray): score = 0 words = sentence.split(" ") for word in words: # 清理单词:转小写+去除前后空白 cleaned_word = word.lower().strip() # 每个单词匹配前重置字典指针到开头 dictionary.seek(0) for entry in dictionary: # 清理字典条目:转小写+去除前后空白 entry_word = entry.strip().lower() if cleaned_word == entry_word: score += 1 break # 找到匹配后立即跳出,减少无效遍历 scores[idx] = score # 将得分对应到正确的数组索引 # 找到得分最高的句子索引 max_score_idx = scores.index(max(scores)) print(scores) print("The sentence is likely to be", sentenceArray[max_score_idx])
方案2:预加载字典到集合(更高效,推荐)
将字典内容一次性加载到集合中,集合的查找是O(1)时间复杂度,既避免指针问题,又大幅提升效率:
def autoDecode(userString): scores = [0] * 25 sentenceArray = decode(userString) # 预加载字典到集合,统一清理格式 with open('engmix.txt') as dictionary: word_set = {entry.strip().lower() for entry in dictionary} for idx, sentence in enumerate(sentenceArray): score = 0 words = sentence.split(" ") for word in words: cleaned_word = word.lower().strip() if cleaned_word in word_set: score += 1 scores[idx] = score max_score_idx = scores.index(max(scores)) print(scores) print("The sentence is likely to be", sentenceArray[max_score_idx])
关键修复点
- 每个单词匹配前重置字典指针(方案1),或预加载字典到集合(方案2),确保每个单词都能正确匹配字典条目。
- 使用
enumerate获取句子索引,将得分对应到scores数组的正确位置。 - 统一处理单词和字典条目的大小写、空白字符,避免格式差异导致匹配失败。
- 最终根据最高得分定位最可能的解密句子,而非固定索引。
内容的提问来源于stack exchange,提问作者user1480135
相关产品推荐
相关产品推荐

