Jupyter调用asr-evaluation计算WER结果与终端不符问题排查
ASR批量计算WER结果异常排查
问题描述
需基于自动语音识别(Automatic Speech Recognition, ASR)转写结果与真实标注(ground truth)转写文本计算词错误率(Word Error Rate, WER)。经测试,GitHub开源仓库asr-evaluation在终端环境下运行可输出正确的WER计算结果。
因业务需批量处理多级目录下的多组转写文件对,在Jupyter Notebook中编写代码导入asr_evaluation模块,遍历目录逐组计算WER时遇到问题:代码可正常执行,但输出指标错误,与同组文件在终端运行得到的正确结果存在明显差异。
无法确定问题源于argparse参数解析传参逻辑,还是asr_evaluation模块内部变量处理机制,需明确调试思路与修复方案。
已提供的相关信息
- 测试文件目录结构:
|-- ASR |--1 |-- ref_groundTruth.txt |-- hyp_fakeASR.txt |--2 |-- ref_groundTruth.txt |-- hyp_groundTruth.txt
- 因模块需接收argparse解析生成的对象作为入参,已从仓库源码复制完整
create_parser函数实现用于生成参数解析器。 - 已提供终端运行、Jupyter运行的结果对比截图,二者同文件计算结果存在明显差异。
- 测试文本内容:
ref_groundTruth.txt、hyp_groundTruth.txt内容一致:
Esto es un texto de prueba para utilizar la libreria ASR, luego de validar, paso al siguiente nivel
hyp_fakeASR.txt内容:
Esto es un text para test para utilizar la libreria ASR, luego de validar, paso al siguiente nivel
- 已编写的Jupyter代码:
# Import libraries from asr_evaluation.asr_evaluation import * import argparse from os import walk # Let's define the parse structure to invoke every time we need past to the module asr_evaluation def create_parser(): parser = argparse.ArgumentParser(description='Evaluate an ASR transcript against a reference transcript.') parser.add_argument('ref', type=argparse.FileType('r'), help='Reference transcript filename') parser.add_argument('hyp', type=argparse.FileType('r'), help='ASR hypothesis filename') print_args = parser.add_mutually_exclusive_group() print_args.add_argument('-i', '--print-instances', action='store_true', help='Print all individual sentences and their errors.') print_args.add_argument('-r', '--print-errors', action='store_true', help='Print all individual sentences that contain errors.') parser.add_argument('--head-ids', action='store_true', help='Hypothesis and reference files have ids in the first token? (Kaldi format)') parser.add_argument('-id', '--tail-ids', '--has-ids', action='store_true', help='Hypothesis and reference files have ids in the last token? (Sphinx format)') parser.add_argument('-c', '--confusions', action='store_true', help='Print tables of which words were confused.') parser.add_argument('-p', '--print-wer-vs-length', action='store_true', help='Print table of average WER grouped by reference sentence length.') parser.add_argument('-m', '--min-word-count', type=int, default=1, metavar='count', help='Minimum word count to show a word in confusions (default 1).') parser.add_argument('-a', '--case-insensitive', action='store_true', help='Down-case the text before running the evaluation.') parser.add_argument('-e', '--remove-empty-refs', action='store_true', help='Skip over any examples where the reference is empty.') return parser # My path is the folder in which the transcripts ASR are and the ground truth mypath = './ASR' #Let's define some macro variables cnt = 1 prefix_ref = 'ref_' prefix_hyp = 'hyp_' # With OS walk we are going to find all the folders related to the transcripted audio which containt \ # the ASR and ground Truth text files for (dirpath, dirnames, filenames) in walk(mypath): print(f'-------{cnt}------') if len(filenames) > 1: groundTruth = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_hyp)]) fakeASR = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_ref)]) parser = create_parser() args = parser.parse_args([groundTruth, fakeASR]) main(args) cnt+=1
问题根因
代码中参考标注文件与ASR假设文件的传参顺序完全写反,是结果异常的核心原因:
- argparse定义的位置参数顺序为:第一个位置传入
ref(参考标注,即ground truth),第二个位置传入hyp(待评估的ASR转写结果)。 - 文件匹配逻辑错误:
groundTruth变量实际取的是hyp_前缀开头的ASR结果文件,fakeASR变量实际取的是ref_前缀开头的真实标注文件。 - 传参时将存了ASR文件路径的
groundTruth放在ref参数位,存了标注文件路径的fakeASR放在hyp参数位,相当于颠倒参考和假设文本计算WER,结果自然与终端正确运行的结果不符。
通用调试思路
遇到模块调用结果与命令行运行不一致的问题,按以下顺序排查即可:
- 核对参数映射:逐行对比命令行运行时的参数顺序、参数值、可选开关配置,与代码中传给argparse的参数列表是否完全一致,重点关注位置参数的顺序。
- 核对文件匹配逻辑:调用核心计算函数前,打印传入的文件路径,读取文件前几行内容确认文件身份,避免文件名前缀匹配写反、路径拼接错误这类低级问题。
- 核对全局状态:如果循环多次调用工具模块,需确认模块内部是否存在未重置的全局计数器、缓存变量,避免前一次计算结果累加到后一次统计中(本问题不存在该情况,但属于批量调用工具类脚本的常见坑)。
修复方案
修正文件匹配与传参逻辑即可,将循环内的代码段修改为:
for (dirpath, dirnames, filenames) in walk(mypath): print(f'-------{cnt}------') if len(filenames) > 1: # 按前缀正确匹配参考文件、ASR结果文件 ref_path = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_ref)]) hyp_path = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_hyp)]) parser = create_parser() # 严格按照ref在前、hyp在后的顺序传参 args = parser.parse_args([ref_path, hyp_path]) main(args) cnt+=1
修改后运行结果将与终端直接调用命令的结果完全一致。
内容的提问来源于stack exchange,提问作者Ale Sosa
相关产品推荐
相关产品推荐

