如何简化Python嵌套循环与判断?优化5x5字谜生成代码效率
优化5x5字谜匹配脚本:简化嵌套循环并提升效率
我是一名业余程序员,编写了一个Python脚本,用于从包含10000个5字母单词的数据集里收集匹配的5x5字谜。该脚本可正常运行,但耗时极长,预计需一年才能完成。我想了解如何简化代码中的嵌套for循环与连续if条件判断,同时提升代码运行效率。
原始代码
allWords = open('allWords.txt').read().splitlines() f = open('finalSolutions.txt', 'a') for v1 in allWords: for h1 in allWords: if h1[0] == v1[0] and h1 != v1: for v2 in allWords: if v2[0] == h1[1] and v2 != h1: for h2 in allWords: if h2[0] == v1[1] and h2[1] == v2[1] and h2 != v1 and h2 != v2: for v3 in allWords: if v3[0] == h1[2] and v3[1] == h2[2] and v3 != h1 and v3 != h2: for h3 in allWords: if h3[0] == v1[2] and h3[1] == v2[2] and h3[2] == v3[2] and h3 != v1 and h3 != v2 and h3 != v3: for v4 in allWords: if v4[0] == h1[3] and v4[1] == h2[3] and v4[2] == h3[3] and v4 != h1 and v4 != h2 and v4 != h3: for h4 in allWords: if h4[0] == v1[3] and h4[1] == v2[3] and h4[2] == v3[3] and h4[3] == v4[3] and h4 != v1 and h4 != v2 and h4 != v3 and h4 != v4: for v5 in allWords: if v5[0] == h1[4] and v5[1] == h2[4] and v5[2] == h3[4] and v5[3] == h4[4] and v5 != h1 and v5 != h2 and v5 != h3 and v5 != h4: for h5 in allWords: if h5[0] == v1[4] and h5[1] == v2[4] and h5[2] == v3[4] and h5[3] == v4[4] and h5[4] == v5[4] and h5 != v1 and h5 != v2 and h5 != v3 and h5 != v4 and h5 != v5: lines = [h1, ', ', h2, ', ', h3, ', ', h4, ', ', h5, ', ', v1, ', ', v2, ', ', v3, ', ', v4, ', ', v5, '\n'] f.writelines(lines) f.close()
输入示例(allWords.txt)
abcde acegi bcdef cdefg defgh efghi opdka bdfhj cegik dfhjl egikm ghijk ijklm
输出示例(finalSolutions.txt)
acegi, bdfhj, cegik, dfhjl, egikm, abcde, cdefg, efghi, ghijk, ijklm abcde, cdefg, efghi, ghijk, ijklm, acegi, bdfhj, cegik, dfhjl, egikm
优化方案
核心思路
原始代码的11层嵌套循环会导致遍历次数达到10000^11量级,完全不可行。优化核心是预构建索引减少无效遍历,同时简化嵌套层级,提前过滤不符合条件的候选。
优化后的代码
def main(): # 读取单词并预处理 with open('allWords.txt', 'r') as f: all_words = f.read().splitlines() word_set = set(all_words) print(f"Loaded {len(all_words)} words") # 预构建各类索引,快速定位符合条件的单词 # 1. 按单词首字符分组 start_with = {c: [w for w in all_words if w[0] == c] for c in 'abcdefghijklmnopqrstuvwxyz'} # 2. 按前n位字符组合映射到单词(直接定位唯一符合条件的单词) pos0_1 = {(w[0], w[1]): w for w in all_words} pos0_1_2 = {(w[0], w[1], w[2]): w for w in all_words} pos0_1_2_3 = {(w[0], w[1], w[2], w[3]): w for w in all_words} with open('finalSolutions.txt', 'a') as out_f: # 逐层构建候选,每一步只处理符合当前条件的单词 for v1 in all_words: # 找h1:首字符和v1相同,且不是v1 for h1 in start_with.get(v1[0], []): if h1 == v1: continue # 找v2:首字符是h1的第2位,且不是h1 for v2 in start_with.get(h1[1], []): if v2 == h1: continue # 直接通过双字符索引找h2,跳过全量遍历 h2 = pos0_1.get((v1[1], v2[1])) if not h2 or h2 in {v1, v2}: continue # 找v3:前两位匹配h1[2]和h2[2] v3 = pos0_1.get((h1[2], h2[2])) if not v3 or v3 in {h1, h2}: continue # 找h3:前三位匹配v1[2]、v2[2]、v3[2] h3 = pos0_1_2.get((v1[2], v2[2], v3[2])) if not h3 or h3 in {v1, v2, v3}: continue # 找v4:前三位匹配h1[3]、h2[3]、h3[3] v4 = pos0_1_2.get((h1[3], h2[3], h3[3])) if not v4 or v4 in {h1, h2, h3}: continue # 找h4:前四位匹配v1[3]、v2[3]、v3[3]、v4[3] h4 = pos0_1_2_3.get((v1[3], v2[3], v3[3], v4[3])) if not h4 or h4 in {v1, v2, v3, v4}: continue # 找v5:前四位匹配h1[4]、h2[4]、h3[4]、h4[4] v5 = pos0_1_2_3.get((h1[4], h2[4], h3[4], h4[4])) if not v5 or v5 in {h1, h2, h3, h4}: continue # 生成h5候选并判断是否在单词集合中 h5_candidate = v1[4] + v2[4] + v3[4] + v4[4] + v5[4] if h5_candidate in word_set and h5_candidate not in {v1, v2, v3, v4, v5}: # 格式化并写入结果 line = f"{h1}, {h2}, {h3}, {h4}, {h5_candidate}, {v1}, {v2}, {v3}, {v4}, {v5}\n" out_f.write(line) if __name__ == "__main__": main()
额外加速建议
- 多核并行:用
multiprocessing模块将遍历v1的任务拆分到多个进程,利用CPU多核能力进一步提速 - 批量写入:积累一定数量的结果后再批量写入文件,减少磁盘IO次数
- 数据去重:提前对输入单词去重,减少无效的重复遍历
内容的提问来源于stack exchange,提问作者Qustom
相关产品推荐
相关产品推荐

