You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure实时说话人分离服务处理UDP/HTTP音频流遇阻求助

Azure Speech Service 实时音频转写与说话人分离问题排查

一、PullAudioStreamCallback 未触发 read 方法的问题

核心原因与修正点

  1. 音频格式不匹配:Azure Speech Service仅支持16kHz采样率、16位位深、单声道PCM格式,若UDP流格式不符,服务不会触发读取逻辑。
  2. 回调类实现错误:必须严格继承PullAudioStreamCallback,且read方法签名、返回值逻辑需符合要求——返回实际读取的字节数,不能返回0(返回0会被判定为流结束)。
  3. 流关联配置遗漏:初始化Speech Recognizer时,需确保自定义Pull流已正确关联到音频配置。

修正后的示例代码

import azure.cognitiveservices.speech as speechsdk
import socket
import time

class UDPAudioPullCallback(speechsdk.audio.PullAudioInputStreamCallback):
    def __init__(self, udp_socket):
        super().__init__()
        self.udp_socket = udp_socket
        self.CHUNK_SIZE = 3200  # 16kHz 16bit单声道,对应100ms音频数据量

    def read(self, buffer: memoryview) -> int:
        # 从UDP读取数据,填充buffer
        data, _ = self.udp_socket.recvfrom(self.CHUNK_SIZE)
        bytes_to_copy = min(len(data), len(buffer))
        buffer[:bytes_to_copy] = data[:bytes_to_copy]
        return bytes_to_copy  # 返回实际读取字节数,不能返回0

# 初始化UDP socket
udp_socket = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
udp_socket.bind(('0.0.0.0', 12345))

# 配置符合要求的音频格式
audio_format = speechsdk.audio.AudioStreamFormat(
    speechsdk.audio.AudioContainerFormat.RAW,
    16000,
    16,
    1
)

# 创建Pull流与音频配置
pull_stream = speechsdk.audio.PullAudioInputStream(callback=UDPAudioPullCallback(udp_socket), format=audio_format)
audio_config = speechsdk.audio.AudioConfig(stream=pull_stream)

# 初始化带说话人分离的识别器
speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域")
speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_SpeakerDiarization_Enabled, "true")
speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_EndSilenceTimeoutMs, "1000")

recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)

# 注册识别回调
def on_recognized(evt):
    if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print(f"转写内容: {evt.result.text}")
        if hasattr(evt.result, 'speaker_id'):
            print(f"说话人ID: {evt.result.speaker_id}")

recognizer.recognized.connect(on_recognized)
recognizer.start_continuous_recognition_async()

# 保持进程运行
while True:
    time.sleep(1)

二、PushAudioInputStream 数据送达但识别结果为空的问题

核心原因与修正点

  1. 音频格式/数据错误:需确保推送的是原始16kHz 16位单声道PCM数据,若为WAV流需先剥离文件头;同时检查字节序为小端序。
  2. 推送速率异常:需模拟实时音频速率推送(如每100ms推送3200字节),过快或过慢都会导致服务无法正确解析。
  3. 说话人分离配置未生效:需确认SpeechServiceConnection_SpeakerDiarization_Enabled已设为"true",且服务区域支持该功能。

修正后的示例代码

import azure.cognitiveservices.speech as speechsdk
import requests
import threading
import time

def push_http_audio(url, push_stream):
    # 从HTTP流读取原始PCM数据
    response = requests.get(url, stream=True)
    CHUNK_SIZE = 3200  # 对应100ms音频数据量
    for chunk in response.iter_content(chunk_size=CHUNK_SIZE):
        if chunk:
            push_stream.write(chunk)
            time.sleep(0.1)  # 模拟实时推送速率
    push_stream.close()

# 配置音频格式
audio_format = speechsdk.audio.AudioStreamFormat(
    speechsdk.audio.AudioContainerFormat.RAW,
    16000,
    16,
    1
)

# 创建Push流与音频配置
push_stream = speechsdk.audio.PushAudioInputStream(format=audio_format)
audio_config = speechsdk.audio.AudioConfig(stream=push_stream)

# 初始化识别器
speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域")
speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_SpeakerDiarization_Enabled, "true")
speech_config.set_property(speechsdk.PropertyId.Speech_LogFilename, "speech_debug.log")  # 开启日志排查

recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config)

# 注册回调,包含无匹配结果的处理
def on_recognized(evt):
    if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech:
        print(f"转写内容: {evt.result.text}")
        if hasattr(evt.result, 'speaker_id'):
            print(f"说话人ID: {evt.result.speaker_id}")
    elif evt.result.reason == speechsdk.ResultReason.NoMatch:
        print(f"无匹配原因: {evt.result.no_match_details}")

recognizer.recognized.connect(on_recognized)

# 启动推送线程
threading.Thread(target=push_http_audio, args=("你的HTTP音频流地址", push_stream)).start()

recognizer.start_continuous_recognition_async()

# 保持进程运行
while True:
    time.sleep(1)

通用排查步骤

  1. 验证音频格式:用ffmpeg将UDP/HTTP流保存为文件,检查格式是否符合要求:
    ffmpeg -i input_stream.raw -ar 16000 -ac 1 -f s16le verified.raw
    
    用麦克风测试流程播放该文件,确认音频本身可被识别。
  2. 升级SDK版本:确保使用最新版Azure Speech SDK:
    pip install --upgrade azure-cognitiveservices-speech
    
  3. 关闭说话人分离测试:先禁用说话人分离功能,确认基础转写正常,再逐步排查分离功能配置。

内容的提问来源于stack exchange,提问作者natbb06

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 21:25:19