You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Wordle四起始词筛选代码?求组合优化学习方向

Wordle四起始词组合筛选的性能优化问题

我已经写了Python代码能筛选出Wordle的最佳单起始词,但我想找最优的四起始词组合(目前用的是spoke、craft、dumpy、light)。现在筛选单起始词只需要几分之一秒,但筛选双起始词要半小时,性能瓶颈明显。想问下该从哪些方向学习这类优化技术?

第一段尝试代码

'''words_file = open("words.txt", 'r')
words_list = words_file.readlines()
words_file.close()'''

words_list = [line.rstrip() for line in open('words.txt', 'r')]

# print(words_list[0])

'''words_5 = []
for word in words_list:
    if len(word) == 5:
        words_5.append(word)'''

words_5 = [word for word in words_list if len(word) == 5]

# print(words_5[0:5])

squished = ''
for word in words_5:
    squished += word

# print(squished[0:20])

frequency = {'a':0, 'b':0, 'c':0, 'd':0, 'e':0, 'f':0, 'g':0, 'h':0, 'i':0, 'j':0, 'k':0, 'l':0, 
    'm':0, 'n':0, 'o':0, 'p':0, 'q':0, 'r':0, 's':0, 't':0, 'u':0, 'v':0, 'w':0, 'x':0, 'y':0, 'z':0, }
for letter in frequency:
    # print(letter)
    frequency[letter] += squished.count(letter)


betterfrequncy = {}
for letter in 'abcdefghijklmnopqrstuvwxyz':
    for position in '12345':
        key = letter + position
        betterfrequncy[key] = len([word for word in words_5 if word[int(position) - 1] == letter])

#print(betterfrequncy)

#print(frequency)

def calcScore(fourWords, penalty):
    score = 0
    squish = ''
    for word in fourWords:
        squish += word
    for letter in frequency :
        score += squish[letter]
        if squish.count(letter) > 1:
            score += penalty * squish.count(letter)
    return score

def betterScore(words):
    score = 0
    squish = ''
    for word in words:
        word = word.lower()
        squish += word
    for word in words:
        word = word.lower()
        #print(word)
        for pos in range(5):
            thisKey = word[pos] + str(pos + 1)
            #print(word[pos])
            score += (squish.count(word[pos]) ** -2) * betterfrequncy[thisKey]

    return score

#print(betterScore(['spoke', 'craft', 'dumpy', 'light']))

print('Crane')
print(betterScore(['Crane']))

top_score = 0
top_word = ''
for word in words_5:
    if top_score < betterScore([word]):
        top_word = word
        top_score = betterScore([word])

print(top_word)
print(top_score)

第二次尝试代码

words_list = [line.rstrip() for line in open('words.txt', 'r')]

words_5 = [word for word in words_list if len(word) == 5]

betterfrequncy = {}
for letter in 'abcdefghijklmnopqrstuvwxyz':
    for position in '12345':
        key = letter + position
        betterfrequncy[key] = len([word for word in words_5 if word[int(position) - 1] == letter])


def betterScore(words):
    score = 0
    squish = ''
    for word in words:
        word = word.lower()
        squish += word
    for word in words:
        word = word.lower()
        #print(word)
        for pos in range(5):
            thisKey = word[pos] + str(pos + 1)
            #print(word[pos])
            score += (squish.count(word[pos]) ** -2) * betterfrequncy[thisKey]

    return score

run = 0
top_score = 0
top_words = []
for word1 in words_5:
    run += 1 
    print("{:.3f}%".format((run / len(words_5)) * 100))
    for word2 in words_5:
        test_words = [word1, word2]
        if top_score < betterScore(test_words):
            top_words = test_words
            top_score = betterScore(test_words)
            
print(top_words)
print(top_score)

优化方向与学习路径

1. 算法层面优化

  • 减少重复计算:betterScore里反复调用squish.count()是低效核心,改用collections.Counter提前统计组合中字母出现次数,后续直接查字典取值,避免多次遍历字符串。
  • 剪枝搜索:先筛选出得分前N的单字词(比如前200个),只在这些候选中组合,大幅减少遍历总量;或者在循环中判断当前组合的理论最高得分是否低于现有最高分,直接跳过后续计算。
  • 避免重复组合:[word1, word2]和[word2, word1]得分完全相同,把内层循环改成for j in range(i+1, len(words_5)),直接砍掉一半计算量。

2. 代码细节优化

  • 预存单字词得分:先遍历所有单字词,把每个词的得分存在字典里,组合时直接累加,不用重复调用betterScore计算单字词得分。
  • 用高效数据结构:比如betterfrequncy的计算可以用生成器替代列表推导式,减少内存占用;或者用NumPy数组存储位置频率,提升计算速度。
  • 减少函数调用开销:把betterScore里的重复逻辑(比如转小写)提前处理,避免在循环内反复执行。

3. 并行计算

  • 用multiprocessing.Pool把单字词列表分成多个子任务,让多核CPU同时处理不同的组合块,直接提升计算速度。注意要避免共享内存的问题,尽量让每个进程独立处理自己的任务。

学习资源方向

  • 算法复杂度基础:先搞懂时间/空间复杂度分析,明白双字词是O(n²)、四字词是O(n⁴)的指数级增长逻辑,从根源理解性能瓶颈。
  • Python性能优化:学习Python内置高效工具(Counter、itertools)、循环优化技巧、避免不必要的对象创建等。
  • 搜索优化算法:深入了解剪枝、分支定界、启发式搜索等方法,这类方法在组合优化问题中能大幅减少计算量。
  • 并行编程:掌握Python多进程、多线程的适用场景,理解GIL对Python并行的影响,学会用multiprocessing或concurrent.futures实现并行计算。

内容的提问来源于stack exchange,提问作者D7G0N _

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 15:20:36