You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Microsoft Cognitive Speech Python SDK集成FastAPI卡顿问题求助

问题描述

我正尝试集成Microsoft语音服务SDK,流程是前端将audioData文件上传至FastAPI后端,再通过SDK发送至微软端点进行发音评估。但每次操作都会卡顿,10秒后出现如下错误:

Speech Recognition canceled: CancellationReason.Error
Error details: Timeout: no recognition result received SessionId: ce96699331684cb7ab4dbb1f619bff10
Info: on_underlying_io_bytes_received: Close frame received
Info: on_underlying_io_bytes_received: closing underlying io.
Info: on_underlying_io_close_complete: uws_state: 6.

我怀疑这个错误和类似的SpeechRecognizer卡顿问题有关,但有两点差异:

  • 我用的Python SDK没有FromWavFileInput方法
  • 尝试添加100KB空缓冲后问题仍未解决

已在Jupyter Notebook中用本地WAV文件测试SDK代码,运行正常,所以问题出在FastAPI集成环节。求解决建议,另外想问是否可以通过API而非SDK进行发音评估?

后端代码

@app.post('/transcribe')
async def transcriptions(audioData: UploadFile = File(...),
                         language: Optional[str] = Form(None)):
    # Read the file content
    audio_content = await audioData.read()
    # Append an empty buffer to the audio content
    audio_content += b'\x00' * 102400  # Add 100KB of silence

    with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio:
        temp_audio.write(audio_content)
        temp_audio_path = temp_audio.name

    # Creates an instance of a speech config with specified subscription key and service region.
    # Replace with your own subscription key and service region (e.g., "westus").
    # Note: The sample is for en-US language.
    print(temp_audio_path)
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
    audio_config = speechsdk.audio.AudioConfig(filename=temp_audio_path)

    reference_text = "I am a boy"
    # Create pronunciation assessment config with json string (JSON format is not recommended)
    enable_miscue, enable_prosody = False, False
    config_json = {
        "GradingSystem": "HundredMark",
        "Granularity": "Phoneme",
        "Dimension": "Comprehensive",
        "ScenarioId": "",  # "" is the default scenario or ask product team for a customized one
        "EnableMiscue": enable_miscue,
        "EnableProsodyAssessment": enable_prosody,
        "NBestPhonemeCount": 0,  # > 0 to enable "spoken phoneme" mode, 0 to disable
    }
    pronunciation_config = speechsdk.PronunciationAssessmentConfig(json_string=json.dumps(config_json))
    pronunciation_config.reference_text = reference_text

    # Create a speech recognizer using a file as audio input.
    language = 'en-US'
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, language=language, audio_config=audio_config)
    # Apply pronunciation assessment config to speech recognizer
    pronunciation_config.apply_to(speech_recognizer)

    result = speech_recognizer.recognize_once_async().get()

    # Check the result
    if result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print('pronunciation assessment for: {}'.format(result.text))
        pronunciation_result = json.loads(result.properties.get(speechsdk.PropertyId.SpeechServiceResponse_JsonResult))
        print('assessment results:\n{}'.format(json.dumps(pronunciation_result, indent=4)))
    elif result.reason == speechsdk.ResultReason.NoMatch:
        print("No speech could be recognized")
    elif result.reason == speechsdk.ResultReason.Canceled:
        cancellation_details = result.cancellation_details
        print("Speech Recognition canceled: {}".format(cancellation_details.reason))
        if cancellation_details.reason == speechsdk.CancellationReason.Error:
            print("Error details: {}".format(cancellation_details.error_details))

    # ignore this - i'm actually doing transcription with this function
    return{'is_subject': True, 'transcription': 'I am a boy'}

if __name__ == "__main__":
    uvicorn.run(app, host="0.0.0.0", port=8080)

前端代码

const filename = new Date().toISOString();
  const formData = new FormData();
  formData.append('audioData', blob, filename);

  axios
      .post('http://localhost:8080/transcribe', formData, {
        'Content-Type': 'multipart/form-data',
      })
      .then((response) => {
        console.log(response.data);
      });

解决建议

针对FastAPI集成的问题排查

  1. 验证上传音频的完整性和格式

    • 前端上传的Blob可能存在格式问题,比如缺少标准WAV文件头、采样率/比特率不匹配微软语音服务要求。可以在后端保存临时文件后,用ffmpeg检查合法性:
      ffmpeg -i temp_audio.wav
      
    • 确保音频是微软支持的格式:推荐16kHz采样率、16位单声道PCM编码的WAV文件。
  2. 优化临时文件处理

    • 当前临时文件写入后未确保内容刷入磁盘,可能导致SDK读取异常。修改代码确保文件写入完整:
      with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio:
          temp_audio.write(audio_content)
          temp_audio.flush()  # 强制将内容写入磁盘
      
    • 处理完后记得删除临时文件,避免磁盘占用:
      import os
      os.unlink(temp_audio_path)
      
  3. 避免阻塞异步函数

    • 在FastAPI的async函数中调用recognize_once_async().get()是同步阻塞操作,会阻塞事件循环导致超时。建议将SDK调用放入线程池:
      from concurrent.futures import ThreadPoolExecutor
      import asyncio
      
      executor = ThreadPoolExecutor()
      
      async def transcriptions(audioData: UploadFile = File(...), language: Optional[str] = Form(None)):
          # ... 前面的代码 ...
          # 把SDK调用放到线程池执行
          result = await asyncio.get_event_loop().run_in_executor(
              executor,
              lambda: speech_recognizer.recognize_once_async().get()
          )
          # ... 后面的代码 ...
      
  4. 修复音频缓冲的添加方式

    • 直接追加空字节会破坏WAV文件头的大小信息,导致SDK无法解析。用wave模块正确修改文件:
      import wave
      from io import BytesIO
      
      audio_io = BytesIO(audio_content)
      with wave.open(audio_io, 'rb') as wav_in:
          params = wav_in.getparams()
          # 计算要追加的空帧数(按16位单声道,每帧2字节)
          add_frames = 102400 // 2
          new_frames = params.nframes + add_frames
          new_params = params._replace(nframes=new_frames)
      
          with tempfile.NamedTemporaryFile(delete=False, suffix='.wav') as temp_audio:
              with wave.open(temp_audio, 'wb') as wav_out:
                  wav_out.setparams(new_params)
                  wav_out.writeframes(wav_in.readframes(params.nframes))
                  # 写入空帧
                  wav_out.writeframes(b'\x00' * add_frames * params.sampwidth * params.nchannels)
      

关于API替代SDK的发音评估

可以直接调用微软语音服务的REST API进行发音评估,无需依赖SDK。核心步骤:

  1. 获取访问令牌:用订阅密钥和区域调用令牌服务获取Bearer令牌
  2. 构造评估请求:发送POST请求到对应区域的语音识别端点,携带参考文本、评估配置和音频数据
  3. 解析响应:从返回的JSON中提取发音分数、音节准确度等评估指标

注意:REST API对音频格式、请求头参数要求严格,需遵循官方文档规范。


内容的提问来源于stack exchange,提问作者Dan Tang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 16:24:54