You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Speech Service批量转写PCMU格式WAV文件遇段错误及解压疑问

问题描述

尝试使用Azure Speech Service转写一批PCMU(μ-law)格式的压缩WAV文件,基于官方文档编写了Python代码。少量文件处理正常,但批量处理约50个文件时频繁出现Segmentation fault错误,出错文件不固定。测试部分文件时,有无解压步骤转写结果一致,怀疑微软推荐的压缩音频处理方法是否生效。运行环境为WSL,已尝试用faulthandler记录错误、提升Python栈限制、添加休眠定时器,问题仍未解决。

代码如下:

import azure.cognitiveservices.speech as speechsdk

def azurespeech_transcribe(audio_filename):
    class BinaryFileReaderCallback(speechsdk.audio.PullAudioInputStreamCallback):
        def __init__(self, filename: str):
            super().__init__()
            self._file_h = open(filename, "rb")

        def read(self, buffer: memoryview) -> int:
            try:
                size = buffer.nbytes
                frames = self._file_h.read(size)
                buffer[:len(frames)] = frames
                return len(frames)
            except Exception as ex:
                print('Exception in `read`: {}'.format(ex))
                raise

        def close(self) -> None:
            try:
                self._file_h.close()
            except Exception as ex:
                print('Exception in `close`: {}'.format(ex))
                raise
    compressed_format = speechsdk.audio.AudioStreamFormat(
        compressed_stream_format=speechsdk.AudioStreamContainerFormat.MULAW
    )
    callback = BinaryFileReaderCallback(filename=audio_filename)
    stream = speechsdk.audio.PullAudioInputStream(
        stream_format=compressed_format,
        pull_stream_callback=callback
    )
    speech_config = speechsdk.SpeechConfig(
        subscription="<my_subscription_key>",
        region="<my_region>",
        speech_recognition_language="en-CA"
    )
    audio_config = speechsdk.audio.AudioConfig(stream=stream)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config)
    result = speech_recognizer.recognize_once()
    return result.text

解决方向

1. 显式回收资源

Segmentation fault大概率和底层资源未及时释放有关,建议在函数结束时显式销毁SDK相关对象,避免内存泄漏累积:

def azurespeech_transcribe(audio_filename):
    # ... 原有代码 ...
    result = speech_recognizer.recognize_once()
    # 显式销毁资源,触发GC回收
    del speech_recognizer
    del audio_config
    del stream
    del callback
    return result.text

2. 复用全局配置

避免在每次转写时重复创建SpeechConfig,将其提升到函数外部复用,减少实例创建开销:

# 全局复用SpeechConfig,只初始化一次
speech_config = speechsdk.SpeechConfig(
    subscription="<my_subscription_key>",
    region="<my_region>",
    speech_recognition_language="en-CA"
)

def azurespeech_transcribe(audio_filename):
    # ... 去掉内部的speech_config创建代码 ...
    audio_config = speechsdk.audio.AudioConfig(stream=stream)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config)
    # ... 其余代码 ...

3. 补全音频格式参数

PCMU标准采样率为8kHz,原代码仅指定容器格式,未明确音频参数,可能导致SDK解析异常。尝试显式配置采样率、位深和声道数:

compressed_format = speechsdk.audio.AudioStreamFormat(
    samples_per_second=8000,
    bits_per_sample=16,
    channels=1,
    compressed_stream_format=speechsdk.AudioStreamContainerFormat.MULAW
)

4. 切换文件输入方式

如果自定义Pull流存在稳定性问题,可先用工具将PCMU转码为未压缩WAV,再使用SDK原生的文件读取接口:

import subprocess
import os

# 用ffmpeg转码(需提前安装ffmpeg)
def convert_pcmu_to_wav(input_path, temp_output):
    subprocess.run([
        "ffmpeg", "-i", input_path,
        "-acodec", "pcm_s16le", "-ar", "8000", "-ac", "1",
        temp_output
    ], check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL)

def azurespeech_transcribe(audio_filename):
    temp_wav = "/tmp/transcode_temp.wav"
    convert_pcmu_to_wav(audio_filename, temp_wav)
    audio_config = speechsdk.audio.AudioConfig(filename=temp_wav)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config)
    result = speech_recognizer.recognize_once()
    os.remove(temp_wav)
    return result.text

5. 调整SDK版本

Segmentation fault多由底层C扩展的内存问题导致,尝试升级到最新版SDK,或回退到已知稳定版本:

# 升级到最新版
pip install --upgrade azure-cognitiveservices-speech

# 回退到稳定版本(如1.30.0)
pip install azure-cognitiveservices-speech==1.30.0

内容的提问来源于stack exchange,提问作者Maxime Gélinas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 10:50:28