Microsoft Cognitive Speech Python SDK集成FastAPI卡顿问题求助
我正尝试集成Microsoft语音服务SDK,流程是前端将audioData文件上传至FastAPI后端,再通过SDK发送至微软端点进行发音评估。但每次操作都会卡顿,10秒后出现如下错误:
Speech Recognition canceled: CancellationReason.Error
Error details: Timeout: no recognition result received SessionId: ce96699331684cb7ab4dbb1f619bff10
Info: on_underlying_io_bytes_received: Close frame received
Info: on_underlying_io_bytes_received: closing underlying io.
Info: on_underlying_io_close_complete: uws_state: 6.
我怀疑这个错误和类似的SpeechRecognizer卡顿问题有关,但有两点差异:
- 我用的Python SDK没有
FromWavFileInput方法 - 尝试添加100KB空缓冲后问题仍未解决
已在Jupyter Notebook中用本地WAV文件测试SDK代码,运行正常,所以问题出在FastAPI集成环节。求解决建议,另外想问是否可以通过API而非SDK进行发音评估?
后端代码
@app.post('/transcribe') async def transcriptions(audioData: UploadFile = File(...), language: Optional[str] = Form(None)): # Read the file content audio_content = await audioData.read() # Append an empty buffer to the audio content audio_content += b'\x00' * 102400 # Add 100KB of silence with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio: temp_audio.write(audio_content) temp_audio_path = temp_audio.name # Creates an instance of a speech config with specified subscription key and service region. # Replace with your own subscription key and service region (e.g., "westus"). # Note: The sample is for en-US language. print(temp_audio_path) speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region) audio_config = speechsdk.audio.AudioConfig(filename=temp_audio_path) reference_text = "I am a boy" # Create pronunciation assessment config with json string (JSON format is not recommended) enable_miscue, enable_prosody = False, False config_json = { "GradingSystem": "HundredMark", "Granularity": "Phoneme", "Dimension": "Comprehensive", "ScenarioId": "", # "" is the default scenario or ask product team for a customized one "EnableMiscue": enable_miscue, "EnableProsodyAssessment": enable_prosody, "NBestPhonemeCount": 0, # > 0 to enable "spoken phoneme" mode, 0 to disable } pronunciation_config = speechsdk.PronunciationAssessmentConfig(json_string=json.dumps(config_json)) pronunciation_config.reference_text = reference_text # Create a speech recognizer using a file as audio input. language = 'en-US' speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, language=language, audio_config=audio_config) # Apply pronunciation assessment config to speech recognizer pronunciation_config.apply_to(speech_recognizer) result = speech_recognizer.recognize_once_async().get() # Check the result if result.reason == speechsdk.ResultReason.RecognizedSpeech: print('pronunciation assessment for: {}'.format(result.text)) pronunciation_result = json.loads(result.properties.get(speechsdk.PropertyId.SpeechServiceResponse_JsonResult)) print('assessment results:\n{}'.format(json.dumps(pronunciation_result, indent=4))) elif result.reason == speechsdk.ResultReason.NoMatch: print("No speech could be recognized") elif result.reason == speechsdk.ResultReason.Canceled: cancellation_details = result.cancellation_details print("Speech Recognition canceled: {}".format(cancellation_details.reason)) if cancellation_details.reason == speechsdk.CancellationReason.Error: print("Error details: {}".format(cancellation_details.error_details)) # ignore this - i'm actually doing transcription with this function return{'is_subject': True, 'transcription': 'I am a boy'} if __name__ == "__main__": uvicorn.run(app, host="0.0.0.0", port=8080)
前端代码
const filename = new Date().toISOString(); const formData = new FormData(); formData.append('audioData', blob, filename); axios .post('http://localhost:8080/transcribe', formData, { 'Content-Type': 'multipart/form-data', }) .then((response) => { console.log(response.data); });
针对FastAPI集成的问题排查
验证上传音频的完整性和格式
- 前端上传的Blob可能存在格式问题,比如缺少标准WAV文件头、采样率/比特率不匹配微软语音服务要求。可以在后端保存临时文件后,用
ffmpeg检查合法性:ffmpeg -i temp_audio.wav - 确保音频是微软支持的格式:推荐16kHz采样率、16位单声道PCM编码的WAV文件。
- 前端上传的Blob可能存在格式问题,比如缺少标准WAV文件头、采样率/比特率不匹配微软语音服务要求。可以在后端保存临时文件后,用
优化临时文件处理
- 当前临时文件写入后未确保内容刷入磁盘,可能导致SDK读取异常。修改代码确保文件写入完整:
with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio: temp_audio.write(audio_content) temp_audio.flush() # 强制将内容写入磁盘 - 处理完后记得删除临时文件,避免磁盘占用:
import os os.unlink(temp_audio_path)
- 当前临时文件写入后未确保内容刷入磁盘,可能导致SDK读取异常。修改代码确保文件写入完整:
避免阻塞异步函数
- 在FastAPI的
async函数中调用recognize_once_async().get()是同步阻塞操作,会阻塞事件循环导致超时。建议将SDK调用放入线程池:from concurrent.futures import ThreadPoolExecutor import asyncio executor = ThreadPoolExecutor() async def transcriptions(audioData: UploadFile = File(...), language: Optional[str] = Form(None)): # ... 前面的代码 ... # 把SDK调用放到线程池执行 result = await asyncio.get_event_loop().run_in_executor( executor, lambda: speech_recognizer.recognize_once_async().get() ) # ... 后面的代码 ...
- 在FastAPI的
修复音频缓冲的添加方式
- 直接追加空字节会破坏WAV文件头的大小信息,导致SDK无法解析。用
wave模块正确修改文件:import wave from io import BytesIO audio_io = BytesIO(audio_content) with wave.open(audio_io, 'rb') as wav_in: params = wav_in.getparams() # 计算要追加的空帧数(按16位单声道,每帧2字节) add_frames = 102400 // 2 new_frames = params.nframes + add_frames new_params = params._replace(nframes=new_frames) with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio: with wave.open(temp_audio, 'wb') as wav_out: wav_out.setparams(new_params) wav_out.writeframes(wav_in.readframes(params.nframes)) # 写入空帧 wav_out.writeframes(b'\x00' * add_frames * params.sampwidth * params.nchannels)
- 直接追加空字节会破坏WAV文件头的大小信息,导致SDK无法解析。用
关于API替代SDK的发音评估
可以直接调用微软语音服务的REST API进行发音评估,无需依赖SDK。核心步骤:
- 获取访问令牌:用订阅密钥和区域调用令牌服务获取Bearer令牌
- 构造评估请求:发送POST请求到对应区域的语音识别端点,携带参考文本、评估配置和音频数据
- 解析响应:从返回的JSON中提取发音分数、音节准确度等评估指标
注意:REST API对音频格式、请求头参数要求严格,需遵循官方文档规范。
内容的提问来源于stack exchange,提问作者Dan Tang

