Azure Speech Service批量转写PCMU格式WAV文件遇段错误及解压疑问
问题描述
尝试使用Azure Speech Service转写一批PCMU(μ-law)格式的压缩WAV文件,基于官方文档编写了Python代码。少量文件处理正常,但批量处理约50个文件时频繁出现Segmentation fault错误,出错文件不固定。测试部分文件时,有无解压步骤转写结果一致,怀疑微软推荐的压缩音频处理方法是否生效。运行环境为WSL,已尝试用faulthandler记录错误、提升Python栈限制、添加休眠定时器,问题仍未解决。
代码如下:
import azure.cognitiveservices.speech as speechsdk def azurespeech_transcribe(audio_filename): class BinaryFileReaderCallback(speechsdk.audio.PullAudioInputStreamCallback): def __init__(self, filename: str): super().__init__() self._file_h = open(filename, "rb") def read(self, buffer: memoryview) -> int: try: size = buffer.nbytes frames = self._file_h.read(size) buffer[:len(frames)] = frames return len(frames) except Exception as ex: print('Exception in `read`: {}'.format(ex)) raise def close(self) -> None: try: self._file_h.close() except Exception as ex: print('Exception in `close`: {}'.format(ex)) raise compressed_format = speechsdk.audio.AudioStreamFormat( compressed_stream_format=speechsdk.AudioStreamContainerFormat.MULAW ) callback = BinaryFileReaderCallback(filename=audio_filename) stream = speechsdk.audio.PullAudioInputStream( stream_format=compressed_format, pull_stream_callback=callback ) speech_config = speechsdk.SpeechConfig( subscription="<my_subscription_key>", region="<my_region>", speech_recognition_language="en-CA" ) audio_config = speechsdk.audio.AudioConfig(stream=stream) speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config) result = speech_recognizer.recognize_once() return result.text
解决方向
1. 显式回收资源
Segmentation fault大概率和底层资源未及时释放有关,建议在函数结束时显式销毁SDK相关对象,避免内存泄漏累积:
def azurespeech_transcribe(audio_filename): # ... 原有代码 ... result = speech_recognizer.recognize_once() # 显式销毁资源,触发GC回收 del speech_recognizer del audio_config del stream del callback return result.text
2. 复用全局配置
避免在每次转写时重复创建SpeechConfig,将其提升到函数外部复用,减少实例创建开销:
# 全局复用SpeechConfig,只初始化一次 speech_config = speechsdk.SpeechConfig( subscription="<my_subscription_key>", region="<my_region>", speech_recognition_language="en-CA" ) def azurespeech_transcribe(audio_filename): # ... 去掉内部的speech_config创建代码 ... audio_config = speechsdk.audio.AudioConfig(stream=stream) speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config) # ... 其余代码 ...
3. 补全音频格式参数
PCMU标准采样率为8kHz,原代码仅指定容器格式,未明确音频参数,可能导致SDK解析异常。尝试显式配置采样率、位深和声道数:
compressed_format = speechsdk.audio.AudioStreamFormat( samples_per_second=8000, bits_per_sample=16, channels=1, compressed_stream_format=speechsdk.AudioStreamContainerFormat.MULAW )
4. 切换文件输入方式
如果自定义Pull流存在稳定性问题,可先用工具将PCMU转码为未压缩WAV,再使用SDK原生的文件读取接口:
import subprocess import os # 用ffmpeg转码(需提前安装ffmpeg) def convert_pcmu_to_wav(input_path, temp_output): subprocess.run([ "ffmpeg", "-i", input_path, "-acodec", "pcm_s16le", "-ar", "8000", "-ac", "1", temp_output ], check=True, stdout=subprocess.DEVNULL, stderr=subprocess.DEVNULL) def azurespeech_transcribe(audio_filename): temp_wav = "/tmp/transcode_temp.wav" convert_pcmu_to_wav(audio_filename, temp_wav) audio_config = speechsdk.audio.AudioConfig(filename=temp_wav) speech_recognizer = speechsdk.SpeechRecognizer(speech_config, audio_config) result = speech_recognizer.recognize_once() os.remove(temp_wav) return result.text
5. 调整SDK版本
Segmentation fault多由底层C扩展的内存问题导致,尝试升级到最新版SDK,或回退到已知稳定版本:
# 升级到最新版 pip install --upgrade azure-cognitiveservices-speech # 回退到稳定版本(如1.30.0) pip install azure-cognitiveservices-speech==1.30.0
内容的提问来源于stack exchange,提问作者Maxime Gélinas
相关产品推荐
相关产品推荐

