You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Azure Conversation Transcriber添加暂停与恢复功能?

解决Azure Conversation Transcriber暂停恢复功能的实现问题

问题背景

我们使用Azure Conversation Transcriber实现带说话人区分(diarization)的实时语音转写功能,需要集成暂停恢复(pause_resume)功能,但尝试多种方法均未奏效。Azure仅提供stop_transcribing_async()函数,该函数会完全终止当前会话。现有代码逻辑为收到"inactive"消息时停止转写器,收到"active"消息时重启转写器,但无法正常运行。

现有代码的问题分析

  1. 转写器重启参数缺失:重启时调用create_conversation_transcriber()未传入必要的连接参数,会导致初始化失败。
  2. 会话中断导致说话人区分失效:每次调用stop_transcribing_async()后,当前会话被完全终止,重启后是全新会话,之前的说话人区分上下文丢失,无法保持转写连续性。
  3. 流关闭后的写入错误:关闭push_stream后,重启时虽创建新流,但原有音频缓冲区和队列的处理逻辑未同步调整,易引发数据写入异常。

可行解决方案

由于Azure Conversation Transcriber未提供原生暂停接口,最优方案是控制音频流输入而非终止转写器:暂停时停止向转写器推送音频数据,恢复时继续推送,保持转写器持续运行,既实现暂停效果,又保留会话上下文和说话人区分状态。

修改后的代码示例

async def receive_audio(uuid, path):
    audio_queue = Queue(maxsize=0)
    transcriber_running = False

    try:
        # 初始化转写器和推送流
        conversation_transcriber, push_stream = create_conversation_transcriber(
            CONNECTIONS.connections[uuid]
        )
        conversation_transcriber.start_transcribing_async().get()
        transcriber_running = True
        current_state = "active"
        CONNECTIONS.connections[uuid]["state"] = current_state

        while True:
            websocket = CONNECTIONS.connections[uuid]["websocket"]
            data = await websocket.recv()

            # 处理状态控制消息
            if isinstance(data, str):
                current_state = data
                CONNECTIONS.connections[uuid]["state"] = current_state
                logger.info(f"Updated state to: {current_state}")
                continue

            # 仅在active状态下处理音频数据
            if current_state == "active" and transcriber_running:
                audio_queue.put_nowait(data)
                while not audio_queue.empty():
                    chunk = get_chunk_from_queue(q=audio_queue, chunk_size=4096)
                    CONNECTIONS.connections[uuid]["audio_buffer"] += chunk
                    push_stream.write(chunk)

    except websockets.exceptions.ConnectionClosed as e:
        logger.info("Connection closed")
        logger.info(e)
    except Exception as e:
        logger.error(f"Error in receive_audio: {e}")
    finally:
        if transcriber_running:
            conversation_transcriber.stop_transcribing_async().get()
            push_stream.close()
        await websocket.close(code=1000)

关键调整说明

  • 保留转写器实例:全程维持同一个转写器会话,避免因重启丢失说话人区分上下文。
  • 状态驱动音频推送:仅当状态为"active"时,才将音频数据写入push_stream;"inactive"状态下暂停推送,转写器因无输入不会产生新的转写结果,实现暂停效果。
  • 简化状态逻辑:移除不必要的转写器销毁和重建步骤,减少异常触发点。

备选方案(需会话续接场景)

若业务必须终止转写器(如长时间暂停节省资源),需在重启时手动维护转写上下文:

  1. 停止转写前保存当前转写结果和说话人标识映射。
  2. 重启转写器后,通过ConversationTranscriber的配置参数传入历史上下文(需确认Azure API支持)。
  3. 恢复音频推送时,在转写结果中合并历史数据。

注意:该方案需依赖Azure Conversation Transcriber的会话续接能力,若API不支持则无法实现说话人区分的连续性。

内容的提问来源于stack exchange,提问作者Googler Thiru

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 07:20:06