You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask应用中使用Azure Speech SDK无磁盘处理上传音频流

Flask中Azure Speech SDK无磁盘音频转文本实现

你的核心问题是错误地将文件流直接传给了PushAudioInputStream的stream_format参数——这个参数需要的是AudioStreamFormat对象(描述音频的编码、采样率等信息),而非原始文件流本身。以下是修正后的完整实现:

修正后的代码

def process_audio_files():
    from azure.cognitiveservices.speech.audio import PushAudioInputStream, AudioConfig, AudioStreamFormat
    import azure.cognitiveservices.speech as speechsdk

    file = request.files['audio-file']
    raw_stream = file.stream

    # 配置Speech服务
    speech_config = speechsdk.SpeechConfig(subscription=speech_key, region=service_region)
    
    # 定义音频格式(根据实际上传的音频格式调整,这里以16kHz、16位单声道PCM为例)
    # 如果是WAV文件,也可以用AudioStreamFormat.from_wav_file_input(raw_stream)自动解析格式
    audio_format = AudioStreamFormat(samples_per_second=16000, bits_per_sample=16, channels=1)
    
    # 创建Push音频输入流
    push_stream = PushAudioInputStream(stream_format=audio_format)
    stream_writer = push_stream.create_writer()

    # 分块读取上传的音频流并写入Push流(避免内存溢出)
    chunk_size = 4096
    while True:
        chunk = raw_stream.read(chunk_size)
        if not chunk:
            break
        stream_writer.write(chunk)
    stream_writer.close()

    # 配置音频输入并初始化识别器
    audio_config = AudioConfig(stream=push_stream)
    auto_detect_source_language_config = speechsdk.languageconfig.AutoDetectSourceLanguageConfig(languages=["en-US", "fr-FR", "es-ES"])
    speech_recognizer = speechsdk.SpeechRecognizer(
        speech_config=speech_config,
        auto_detect_source_language_config=auto_detect_source_language_config,
        audio_config=audio_config
    )

    # 执行识别并返回结果
    result = speech_recognizer.recognize_once()
    if result.reason == speechsdk.ResultReason.RecognizedSpeech:
        return {
            "transcript": result.text,
            "detected_language": result.properties[speechsdk.PropertyId.SpeechServiceConnection_AutoDetectSourceLanguageResult]
        }
    elif result.reason == speechsdk.ResultReason.NoMatch:
        return {"error": f"No speech detected: {result.no_match_details}"}
    elif result.reason == speechsdk.ResultReason.Canceled:
        return {"error": f"Recognition canceled: {result.cancellation_details.reason}"}

关键注意事项

  • 音频格式匹配:必须确保AudioStreamFormat的参数与上传音频的实际格式一致(采样率、位深、声道数)。如果上传的是标准WAV文件,可改用AudioStreamFormat.from_wav_file_input(raw_stream)自动解析格式,调用后需重置文件流指针:raw_stream.seek(0)。
  • 分块写入逻辑:采用分块读取写入的方式,避免一次性加载大音频文件导致内存占用过高。
  • 流生命周期管理:写入完成后必须关闭stream_writer,否则识别器无法判断流的结束位置,可能导致识别超时。

内容的提问来源于stack exchange,提问作者HoverCraft

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 20:57:57