Azure实时说话人分离服务处理UDP/HTTP音频流遇阻求助
Azure Speech Service 实时音频转写与说话人分离问题排查
一、PullAudioStreamCallback 未触发 read 方法的问题
核心原因与修正点
- 音频格式不匹配:Azure Speech Service仅支持16kHz采样率、16位位深、单声道PCM格式,若UDP流格式不符,服务不会触发读取逻辑。
- 回调类实现错误:必须严格继承
PullAudioStreamCallback,且read方法签名、返回值逻辑需符合要求——返回实际读取的字节数,不能返回0(返回0会被判定为流结束)。 - 流关联配置遗漏:初始化Speech Recognizer时,需确保自定义Pull流已正确关联到音频配置。
修正后的示例代码
import azure.cognitiveservices.speech as speechsdk import socket import time class UDPAudioPullCallback(speechsdk.audio.PullAudioInputStreamCallback): def __init__(self, udp_socket): super().__init__() self.udp_socket = udp_socket self.CHUNK_SIZE = 3200 # 16kHz 16bit单声道,对应100ms音频数据量 def read(self, buffer: memoryview) -> int: # 从UDP读取数据,填充buffer data, _ = self.udp_socket.recvfrom(self.CHUNK_SIZE) bytes_to_copy = min(len(data), len(buffer)) buffer[:bytes_to_copy] = data[:bytes_to_copy] return bytes_to_copy # 返回实际读取字节数,不能返回0 # 初始化UDP socket udp_socket = socket.socket(socket.AF_INET, socket.SOCK_DGRAM) udp_socket.bind(('0.0.0.0', 12345)) # 配置符合要求的音频格式 audio_format = speechsdk.audio.AudioStreamFormat( speechsdk.audio.AudioContainerFormat.RAW, 16000, 16, 1 ) # 创建Pull流与音频配置 pull_stream = speechsdk.audio.PullAudioInputStream(callback=UDPAudioPullCallback(udp_socket), format=audio_format) audio_config = speechsdk.audio.AudioConfig(stream=pull_stream) # 初始化带说话人分离的识别器 speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域") speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_SpeakerDiarization_Enabled, "true") speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_EndSilenceTimeoutMs, "1000") recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) # 注册识别回调 def on_recognized(evt): if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech: print(f"转写内容: {evt.result.text}") if hasattr(evt.result, 'speaker_id'): print(f"说话人ID: {evt.result.speaker_id}") recognizer.recognized.connect(on_recognized) recognizer.start_continuous_recognition_async() # 保持进程运行 while True: time.sleep(1)
二、PushAudioInputStream 数据送达但识别结果为空的问题
核心原因与修正点
- 音频格式/数据错误:需确保推送的是原始16kHz 16位单声道PCM数据,若为WAV流需先剥离文件头;同时检查字节序为小端序。
- 推送速率异常:需模拟实时音频速率推送(如每100ms推送3200字节),过快或过慢都会导致服务无法正确解析。
- 说话人分离配置未生效:需确认
SpeechServiceConnection_SpeakerDiarization_Enabled已设为"true",且服务区域支持该功能。
修正后的示例代码
import azure.cognitiveservices.speech as speechsdk import requests import threading import time def push_http_audio(url, push_stream): # 从HTTP流读取原始PCM数据 response = requests.get(url, stream=True) CHUNK_SIZE = 3200 # 对应100ms音频数据量 for chunk in response.iter_content(chunk_size=CHUNK_SIZE): if chunk: push_stream.write(chunk) time.sleep(0.1) # 模拟实时推送速率 push_stream.close() # 配置音频格式 audio_format = speechsdk.audio.AudioStreamFormat( speechsdk.audio.AudioContainerFormat.RAW, 16000, 16, 1 ) # 创建Push流与音频配置 push_stream = speechsdk.audio.PushAudioInputStream(format=audio_format) audio_config = speechsdk.audio.AudioConfig(stream=push_stream) # 初始化识别器 speech_config = speechsdk.SpeechConfig(subscription="你的订阅密钥", region="你的服务区域") speech_config.set_property(speechsdk.PropertyId.SpeechServiceConnection_SpeakerDiarization_Enabled, "true") speech_config.set_property(speechsdk.PropertyId.Speech_LogFilename, "speech_debug.log") # 开启日志排查 recognizer = speechsdk.SpeechRecognizer(speech_config=speech_config, audio_config=audio_config) # 注册回调,包含无匹配结果的处理 def on_recognized(evt): if evt.result.reason == speechsdk.ResultReason.RecognizedSpeech: print(f"转写内容: {evt.result.text}") if hasattr(evt.result, 'speaker_id'): print(f"说话人ID: {evt.result.speaker_id}") elif evt.result.reason == speechsdk.ResultReason.NoMatch: print(f"无匹配原因: {evt.result.no_match_details}") recognizer.recognized.connect(on_recognized) # 启动推送线程 threading.Thread(target=push_http_audio, args=("你的HTTP音频流地址", push_stream)).start() recognizer.start_continuous_recognition_async() # 保持进程运行 while True: time.sleep(1)
通用排查步骤
- 验证音频格式:用ffmpeg将UDP/HTTP流保存为文件,检查格式是否符合要求:
用麦克风测试流程播放该文件,确认音频本身可被识别。ffmpeg -i input_stream.raw -ar 16000 -ac 1 -f s16le verified.raw - 升级SDK版本:确保使用最新版Azure Speech SDK:
pip install --upgrade azure-cognitiveservices-speech - 关闭说话人分离测试:先禁用说话人分离功能,确认基础转写正常,再逐步排查分离功能配置。
内容的提问来源于stack exchange,提问作者natbb06
相关产品推荐
相关产品推荐

