You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用SpaCy计算两个目录下匹配转录文件的余弦相似度

双目录匹配遍历计算转录文本相似度解决方案

核心实现逻辑

你的需求是配对两个目录下匹配的文件,计算SpaCy相似度,核心可以用Python内置的zip()函数实现等长列表的逐元素配对遍历,刚好适配你两个目录文件数量一致的场景。

原代码错误原因

第一种写法报错原因

你写的for human_file, api_file in os.listdir(human_directory), os.listdir(api_directory)属于语法逻辑错误:Python会将逗号分隔的两个列表识别为一个二元组进行遍历,第一次迭代就会把第一个目录的完整文件列表尝试赋值给human_file和api_file两个变量,直接触发变量数量不匹配的报错。

第二种写法仅输出2个结果的原因

你写的for i in (0, (len(human_directory) - 1))实际是遍历一个仅包含2个元素的元组(第一个元素是0,第二个是目录文件列表的最后一位索引),所以只会循环两次。如果要按索引遍历所有元素,需要改为for i in range(len(human_directory))。

可行实现代码

你后续更新的使用zip()的写法已经可以正常运行,针对你遇到的UTF-16字符中断问题,可以在读取文件时指定编码规则做兼容,优化后的代码如下:

import os
import spacy

# 加载SpaCy小模型
nlp_small = spacy.load('en_core_web_sm')

# 设置转录文件所在目录
human_dir = os.path.join(os.getcwd(), "00_data/Human Transcripts")
api_dir = os.path.join(os.getcwd(), "00_data/Watson Scripts")

# 配对遍历两个目录下的文件
for human_file, api_file in zip(os.listdir(human_dir), os.listdir(api_dir)):
    # 构造完整文件路径
    human_path = os.path.join(human_dir, human_file)
    api_path = os.path.join(api_dir, api_file)
    # 读取文件时添加编码兼容规则,避免特殊字符中断运行
    with open(human_path, 'r', encoding='utf-8', errors='ignore') as f:
        human_text = f.read()
    with open(api_path, 'r', encoding='utf-8', errors='ignore') as f:
        api_text = f.read()
    # 计算相似度
    human_doc = nlp_small(human_text)
    api_doc = nlp_small(api_text)
    # 打印结果
    print(f"小模型相似度:{human_file} <-> {api_file} 得分:{human_doc.similarity(api_doc):.4f}")

运行效果参考

已跑出的部分结果如下,相似度均在0.87~0.96区间,符合同音频转录的匹配特征:

小模型相似度:human_10.txt <-> watson_10.csv 得分:0.9275
小模型相似度:human_11.txt <-> watson_11.csv 得分:0.9349
小模型相似度:human_12.txt <-> watson_12.csv 得分:0.9362
小模型相似度:human_13.txt <-> watson_13.csv 得分:0.9557
小模型相似度:human_14.txt <-> watson_14.csv 得分:0.9089
小模型相似度:human_15.txt <-> watson_15.csv 得分:0.9479
小模型相似度:human_16.txt <-> watson_16.csv 得分:0.9600
小模型相似度:human_17.txt <-> watson_17.csv 得分:0.9368
小模型相似度:human_18.txt <-> watson_18.csv 得分:0.8761
小模型相似度:human_2.txt <-> watson_2.csv 得分:0.9185
小模型相似度:human_3.txt <-> watson_3.csv 得分:0.9287
小模型相似度:human_4.txt <-> watson_4.csv 得分:0.9416
小模型相似度:human_5.txt <-> watson_5.csv 得分:0.9159
小模型相似度:human_6.txt <-> watson_6.csv 得分:0.9353

后续优化建议

修复编码问题后可以将逻辑封装为通用函数,支持切换大小模型、输出结果到CSV文件等需求,方便批量测试不同API的转录效果。

内容的提问来源于stack exchange,提问作者jtoepp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 03:54:04