You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Speech-Text报错:_io.BytesIO对象无_handle属性求助

解决Azure语音识别中BytesIO流触发的_handle属性错误问题

问题根源

Azure Speech SDK的AudioConfig(stream=...)参数不支持直接传入Python原生的BytesIO对象,它要求使用SDK内置的音频流类型。普通的文件类对象(比如BytesIO)没有SDK底层需要的_handle属性,因此会触发'_io.BytesIO' object has no attribute '_handle'错误。

解决方案

需要将解码后的BytesIO包装为SDK兼容的PullAudioInputStream,具体步骤如下:

修改后的完整代码

import azure.functions as func
import logging
import base64
import io
import azure.cognitiveservices.speech as speechsdk
from azure.cognitiveservices.speech.audio import PullAudioInputStream, AudioStreamFormat

# 替换为你的Azure语音服务密钥和区域
speech_key = "your-speech-key"
service_region = "your-service-region"

def speech_recognize_continuous_from_file(audio_stream):
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
    
    # 使用SDK兼容的音频流初始化AudioConfig
    audio_config = speechsdk.audio.AudioConfig(stream=audio_stream)
    speech_recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)

    # 处理识别结果的回调逻辑
    done = False
    def stop_recognition(evt):
        nonlocal done
        done = True

    # 绑定识别事件监听
    speech_recognizer.recognized.connect(lambda evt: logging.info(f"识别结果: {evt.result.text}"))
    speech_recognizer.session_stopped.connect(stop_recognition)
    speech_recognizer.canceled.connect(stop_recognition)

    # 启动连续识别并等待结束
    speech_recognizer.start_continuous_recognition()
    while not done:
        pass
    speech_recognizer.stop_continuous_recognition()

def transcriptionFunction(req: func.HttpRequest) -> func.HttpResponse:
    logging.info('Python HTTP trigger function processed a request.')

    try:
        req_body = req.get_json()
        audioBase64 = req_body.get('audioBase64')

        # Base64解码为BytesIO
        decodedAudio = base64.b64decode(audioBase64)
        audioIO = io.BytesIO(decodedAudio)

        # 定义音频流格式(需与你的WAV文件参数完全匹配)
        audio_format = AudioStreamFormat(
            samples_per_second=16000,  # 替换为实际采样率
            bits_per_sample=16,         # 替换为实际采样位深
            channels=1                  # 替换为实际声道数
        )
        
        # 定义读取回调,供SDK从BytesIO中拉取音频数据
        def read_callback(size):
            return audioIO.read(size)
        
        # 将BytesIO包装为SDK兼容的PullAudioInputStream
        audio_stream = PullAudioInputStream(audio_format, read_callback)

        # 启动语音识别
        speech_recognize_continuous_from_file(audio_stream)

        return func.HttpResponse("Check Server Console for response", status_code=200)
    except Exception as e:
        logging.error(f"处理失败: {str(e)}")
        return func.HttpResponse(f"Error: {str(e)}", status_code=500)

核心要点

  • PullAudioInputStream:这是SDK专为自定义音频输入场景设计的类,通过读取回调函数实现从BytesIO中按需获取音频数据。
  • 音频流格式匹配:AudioStreamFormat的参数必须与你的WAV文件实际参数完全一致(采样率、位深、声道数),你已确认参数符合要求,直接填入对应数值即可。
  • 回调逻辑:read_callback负责每次向SDK返回指定大小的音频数据,是连接BytesIO和SDK的关键桥梁。

内容的提问来源于stack exchange,提问作者Matt Drinkall

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 11:47:08