You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Jupyter调用asr-evaluation计算WER结果与终端不符问题排查

ASR批量计算WER结果异常排查

问题描述

需基于自动语音识别(Automatic Speech Recognition, ASR)转写结果与真实标注(ground truth)转写文本计算词错误率(Word Error Rate, WER)。经测试,GitHub开源仓库asr-evaluation在终端环境下运行可输出正确的WER计算结果。
因业务需批量处理多级目录下的多组转写文件对,在Jupyter Notebook中编写代码导入asr_evaluation模块,遍历目录逐组计算WER时遇到问题:代码可正常执行,但输出指标错误,与同组文件在终端运行得到的正确结果存在明显差异。
无法确定问题源于argparse参数解析传参逻辑,还是asr_evaluation模块内部变量处理机制,需明确调试思路与修复方案。

已提供的相关信息

  • 测试文件目录结构:
|-- ASR
    |--1
        |-- ref_groundTruth.txt
        |-- hyp_fakeASR.txt
    |--2
        |-- ref_groundTruth.txt
        |-- hyp_groundTruth.txt
  • 因模块需接收argparse解析生成的对象作为入参,已从仓库源码复制完整create_parser函数实现用于生成参数解析器。
  • 已提供终端运行、Jupyter运行的结果对比截图,二者同文件计算结果存在明显差异。
  • 测试文本内容:
    • ref_groundTruth.txt、hyp_groundTruth.txt内容一致:
Esto es un texto de prueba para utilizar la libreria ASR, luego de validar, paso al siguiente nivel
  • hyp_fakeASR.txt内容:
Esto es un text para test para utilizar la libreria ASR, luego de validar, paso al siguiente nivel
  • 已编写的Jupyter代码:
# Import libraries
from asr_evaluation.asr_evaluation import *
import argparse
from os import walk

# Let's define the parse structure to invoke every time we need past to the module asr_evaluation
def create_parser():
    parser = argparse.ArgumentParser(description='Evaluate an ASR transcript against a reference transcript.')
    parser.add_argument('ref', type=argparse.FileType('r'), help='Reference transcript filename')
    parser.add_argument('hyp', type=argparse.FileType('r'), help='ASR hypothesis filename')
    print_args = parser.add_mutually_exclusive_group()
    print_args.add_argument('-i', '--print-instances', action='store_true',
                            help='Print all individual sentences and their errors.')
    print_args.add_argument('-r', '--print-errors', action='store_true',
                            help='Print all individual sentences that contain errors.')
    parser.add_argument('--head-ids', action='store_true',
                        help='Hypothesis and reference files have ids in the first token? (Kaldi format)')
    parser.add_argument('-id', '--tail-ids', '--has-ids', action='store_true',
                        help='Hypothesis and reference files have ids in the last token? (Sphinx format)')
    parser.add_argument('-c', '--confusions', action='store_true', help='Print tables of which words were confused.')
    parser.add_argument('-p', '--print-wer-vs-length', action='store_true',
                        help='Print table of average WER grouped by reference sentence length.')
    parser.add_argument('-m', '--min-word-count', type=int, default=1, metavar='count',
                        help='Minimum word count to show a word in confusions (default 1).')
    parser.add_argument('-a', '--case-insensitive', action='store_true',
                        help='Down-case the text before running the evaluation.')
    parser.add_argument('-e', '--remove-empty-refs', action='store_true',
                        help='Skip over any examples where the reference is empty.')
    
    return parser

# My path is the folder in which the transcripts ASR are and the ground truth
mypath = './ASR'

#Let's define some macro variables
cnt = 1
prefix_ref = 'ref_'
prefix_hyp = 'hyp_'

# With OS walk we are going to find all the folders related to the transcripted audio which containt \
# the ASR and ground Truth text files
for (dirpath, dirnames, filenames) in walk(mypath):
    print(f'-------{cnt}------')

    if len(filenames) > 1:
        groundTruth = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_hyp)])
        fakeASR = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_ref)])

        parser = create_parser()
        args = parser.parse_args([groundTruth, fakeASR])
        main(args)
        
    cnt+=1

问题根因

代码中参考标注文件与ASR假设文件的传参顺序完全写反,是结果异常的核心原因:

  1. argparse定义的位置参数顺序为:第一个位置传入ref(参考标注,即ground truth),第二个位置传入hyp(待评估的ASR转写结果)。
  2. 文件匹配逻辑错误:groundTruth变量实际取的是hyp_前缀开头的ASR结果文件,fakeASR变量实际取的是ref_前缀开头的真实标注文件。
  3. 传参时将存了ASR文件路径的groundTruth放在ref参数位,存了标注文件路径的fakeASR放在hyp参数位,相当于颠倒参考和假设文本计算WER,结果自然与终端正确运行的结果不符。

通用调试思路

遇到模块调用结果与命令行运行不一致的问题,按以下顺序排查即可:

  • 核对参数映射:逐行对比命令行运行时的参数顺序、参数值、可选开关配置,与代码中传给argparse的参数列表是否完全一致,重点关注位置参数的顺序。
  • 核对文件匹配逻辑:调用核心计算函数前,打印传入的文件路径,读取文件前几行内容确认文件身份,避免文件名前缀匹配写反、路径拼接错误这类低级问题。
  • 核对全局状态:如果循环多次调用工具模块,需确认模块内部是否存在未重置的全局计数器、缓存变量,避免前一次计算结果累加到后一次统计中(本问题不存在该情况,但属于批量调用工具类脚本的常见坑)。

修复方案

修正文件匹配与传参逻辑即可,将循环内的代码段修改为:

for (dirpath, dirnames, filenames) in walk(mypath):
    print(f'-------{cnt}------')
    if len(filenames) > 1:
        # 按前缀正确匹配参考文件、ASR结果文件
        ref_path = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_ref)])
        hyp_path = dirpath + '/' + ''.join([word for word in filenames if word.startswith(prefix_hyp)])
        parser = create_parser()
        # 严格按照ref在前、hyp在后的顺序传参
        args = parser.parse_args([ref_path, hyp_path])
        main(args)
    cnt+=1

修改后运行结果将与终端直接调用命令的结果完全一致。


内容的提问来源于stack exchange,提问作者Ale Sosa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 18:39:55