Python使用SpaCy计算两个目录下匹配转录文件的余弦相似度
双目录匹配遍历计算转录文本相似度解决方案
核心实现逻辑
你的需求是配对两个目录下匹配的文件,计算SpaCy相似度,核心可以用Python内置的zip()函数实现等长列表的逐元素配对遍历,刚好适配你两个目录文件数量一致的场景。
原代码错误原因
第一种写法报错原因
你写的for human_file, api_file in os.listdir(human_directory), os.listdir(api_directory)属于语法逻辑错误:Python会将逗号分隔的两个列表识别为一个二元组进行遍历,第一次迭代就会把第一个目录的完整文件列表尝试赋值给human_file和api_file两个变量,直接触发变量数量不匹配的报错。
第二种写法仅输出2个结果的原因
你写的for i in (0, (len(human_directory) - 1))实际是遍历一个仅包含2个元素的元组(第一个元素是0,第二个是目录文件列表的最后一位索引),所以只会循环两次。如果要按索引遍历所有元素,需要改为for i in range(len(human_directory))。
可行实现代码
你后续更新的使用zip()的写法已经可以正常运行,针对你遇到的UTF-16字符中断问题,可以在读取文件时指定编码规则做兼容,优化后的代码如下:
import os import spacy # 加载SpaCy小模型 nlp_small = spacy.load('en_core_web_sm') # 设置转录文件所在目录 human_dir = os.path.join(os.getcwd(), "00_data/Human Transcripts") api_dir = os.path.join(os.getcwd(), "00_data/Watson Scripts") # 配对遍历两个目录下的文件 for human_file, api_file in zip(os.listdir(human_dir), os.listdir(api_dir)): # 构造完整文件路径 human_path = os.path.join(human_dir, human_file) api_path = os.path.join(api_dir, api_file) # 读取文件时添加编码兼容规则,避免特殊字符中断运行 with open(human_path, 'r', encoding='utf-8', errors='ignore') as f: human_text = f.read() with open(api_path, 'r', encoding='utf-8', errors='ignore') as f: api_text = f.read() # 计算相似度 human_doc = nlp_small(human_text) api_doc = nlp_small(api_text) # 打印结果 print(f"小模型相似度:{human_file} <-> {api_file} 得分:{human_doc.similarity(api_doc):.4f}")
运行效果参考
已跑出的部分结果如下,相似度均在0.87~0.96区间,符合同音频转录的匹配特征:
小模型相似度:human_10.txt <-> watson_10.csv 得分:0.9275 小模型相似度:human_11.txt <-> watson_11.csv 得分:0.9349 小模型相似度:human_12.txt <-> watson_12.csv 得分:0.9362 小模型相似度:human_13.txt <-> watson_13.csv 得分:0.9557 小模型相似度:human_14.txt <-> watson_14.csv 得分:0.9089 小模型相似度:human_15.txt <-> watson_15.csv 得分:0.9479 小模型相似度:human_16.txt <-> watson_16.csv 得分:0.9600 小模型相似度:human_17.txt <-> watson_17.csv 得分:0.9368 小模型相似度:human_18.txt <-> watson_18.csv 得分:0.8761 小模型相似度:human_2.txt <-> watson_2.csv 得分:0.9185 小模型相似度:human_3.txt <-> watson_3.csv 得分:0.9287 小模型相似度:human_4.txt <-> watson_4.csv 得分:0.9416 小模型相似度:human_5.txt <-> watson_5.csv 得分:0.9159 小模型相似度:human_6.txt <-> watson_6.csv 得分:0.9353
后续优化建议
修复编码问题后可以将逻辑封装为通用函数,支持切换大小模型、输出结果到CSV文件等需求,方便批量测试不同API的转录效果。
内容的提问来源于stack exchange,提问作者jtoepp
相关产品推荐
相关产品推荐

