You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现单词相似度查询时for循环仅执行一次就停止问题

问题修复方案

问题原因

  • find_closest函数每次执行都会调用create_data()读取sys.stdin流,而标准输入流是单次读取的对象,第一次读取完成后指针就会停在文件末尾,后续再次读取不会返回任何内容。
  • 处理第一个传入参数时数据加载正常,处理第二个参数时create_data()返回空字典,因此没有匹配的相似单词输出。

修复方案

将数据加载逻辑移到参数循环外,全局仅加载一次数据,修改后代码如下:

import sys
from collections import defaultdict
# 此处需保留你原有lev_dist函数的实现逻辑
def create_data():
    data = defaultdict(int)
    value = 0
    for line in sys.stdin:
        [ident, user, text, terms] = line.rstrip().split('\t')
        for word in terms.split():
            data[word] = value
    return data

def find_closest(word, data):
    data_with_distance = defaultdict(int)
    for key in data:
        distance = lev_dist(word, key)
        data_with_distance[key] = distance
    return {k: v for k, v in sorted(data_with_distance.items(), key=lambda item: item[1])}

def main():
    if len(sys.argv) > 1:
        # 程序启动时仅加载一次数据
        data = create_data()
        for w in sys.argv[1:]:
            print(f"\n{w} is close to:\n")
            closest = find_closest(w, data)
            closest_words = [k for k, v in closest.items() if v < 4]
            for close in closest_words:
                print(close, end=", ")
    else:
        sys.stderr.write("no argument\n")

if __name__ == '__main__':
    main()

调整后避免了重复读取输入流的问题,同时降低了不必要的重复计算开销,所有传入的参数都能正常匹配到相似单词。

内容的提问来源于stack exchange,提问作者Keizer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 20:27:00