使用Azure Speech Recognizer时遭遇SPXERR_INVALID_HEADER错误求助
Azure Speech 语音转文本报错 SPXERR_INVALID_HEADER (0xa)
问题描述
使用Azure Speech服务实现语音转文本时,运行Python代码触发RuntimeError,错误码0xa(SPXERR_INVALID_HEADER)。已确认拥有Azure账号及Speech服务的订阅密钥、区域信息,代码及错误栈如下:
代码
import azure.cognitiveservices.speech as speechsdk subscription_key = "密钥内容" service_region = "区域信息" audio_file = "音频文件路径" speech_config = speechsdk.SpeechConfig(subscription=subscription_key, region=service_region) audio_input = speechsdk.AudioConfig(filename=audio_file) speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input, language="tr") result = speech_recognizer.recognize_once() if result.reason == speechsdk.ResultReason.RecognizedSpeech: print("文本: {}".format(result.text)) elif result.reason == speechsdk.ResultReason.NoMatch: print("未匹配到内容: {}".format(result.no_match_details.reason)) elif result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print("识别已取消: {}".format(cancellation_details.reason)) if cancellation_details.reason == speechsdk.CancellationReason.Error: print("错误详情: {}".format(cancellation_details.reason_details))
错误信息
---> 16 speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_input, language="tr") RuntimeError: Exception with error code: [CALL STACK BEGIN] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1aa875) [0x7fd4c79aa875] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1aaf64) [0x7fd4c79aaf64] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1ac537) [0x7fd4c79ac537] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1acbee) [0x7fd4c79acbee] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1919bc) [0x7fd4c79919bc] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x194b3e) [0x7fd4c7994b3e] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1346c5) [0x7fd4c79346c5] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x19b253) [0x7fd4c799b253] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x19b6c2) [0x7fd4c799b6c2] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x13e447) [0x7fd4c793e447] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x1e65f6) [0x7fd4c79e65f6] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x139b9b) [0x7fd4c7939b9b] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(+0x20dfe2) [0x7fd4c7a0dfe2] /usr/local/lib/python3.10/dist-packages/azure/cognitiveservices/speech/libMicrosoft.CognitiveServices.Speech.core.so(recognizer_create_speech_recognizer_from_source_lang_config+0x116) [0x7fd4c78bf641] /lib/x86_64-linux-gnu/libffi.so.8(+0x7e2e) [0x7fd509008e2e] /lib/x86_64-linux-gnu/libffi.so.8(+0x4493) [0x7fd509005493] /usr/lib/python3.10/lib-dynload/_ctypes.cpython-310-x86_64-linux-gnu.so(+0xa3e9) [0x7fd50902e3e9] [CALL STACK END] Exception with an error code: 0xa (SPXERR_INVALID_HEADER)
解决方案
SPXERR_INVALID_HEADER错误核心原因是音频文件格式不兼容或文件损坏,按以下步骤排查修复:
检查音频格式兼容性
Azure Speech服务优先支持WAV(PCM编码,16kHz/8kHz采样率,16位单声道),也支持MP3、OGG等格式。如果当前文件是其他格式,转换为标准WAV后重试。验证音频文件完整性
直接播放音频确认是否能正常播放,若文件损坏、头部信息缺失或截断,会触发该错误。重新获取完整的音频文件后再测试。显式指定音频格式参数
若使用非默认格式的音频,在AudioConfig中明确指定编码、采样率和声道数,示例:audio_format = speechsdk.AudioStreamFormat(encoding=speechsdk.AudioEncoding.MP3, sample_rate=16000, channels=1) audio_input = speechsdk.AudioConfig(filename=audio_file, format=audio_format)参数需与实际音频属性一致。
确认语言代码有效性
代码中language="tr"是土耳其语的正确代码,但需确保你的Speech服务资源支持该语言。可在Azure门户Speech服务页面查看支持的语言列表,若不支持则更换为兼容的语言代码。升级Speech SDK版本
旧版本SDK可能存在格式兼容bug,运行pip install --upgrade azure-cognitiveservices-speech升级到最新版本后重试。
内容的提问来源于stack exchange,提问作者SilentHorse
相关产品推荐
相关产品推荐

