使用Azure Speech Recognition实现Speaker Identity报错咨询
Azure语音识别区分说话人ID报错的解决方法
错误原因
- 普通
SpeechRecognizer不提供说话人识别能力,PropertyId枚举里根本没有Speaker这个选项,你调用的API本身就不支持返回说话人ID。 - 要区分对话中的不同说话人,必须使用Azure语音服务的**对话转录(Conversation Transcription)**功能,这是独立于普通语音识别的专项能力。
解决步骤
1. 配置Azure服务与依赖
- 确保你的Azure语音服务实例已启用,且使用的区域(比如eastus2)支持对话转录功能。
- 安装必要的NuGet包:
- 核心包:
Microsoft.CognitiveServices.Speech(保持1.35.0版本即可) - 对话转录扩展包:
Microsoft.CognitiveServices.Speech.Transcription
- 核心包:
2. 修改代码实现多说话人识别
替换原来的SpeechRecognizer为ConversationTranscriber,它会自动识别对话中的说话人并返回对应的ID。修改后的完整代码如下:
private async void ProcessWavFile(string filePath) { try { string subscriptionKey = "mykey"; string region = "eastus2"; var config = SpeechConfig.FromSubscription(subscriptionKey, region); // 启用对话转录的多说话人识别模式 config.SetProperty(PropertyId.ConversationTranscription_InRoomAndOnline, "true"); using (var audioConfig = AudioConfig.FromWavFileInput(filePath)) using (var transcriber = new ConversationTranscriber(config, audioConfig)) { // 订阅转录完成事件,获取说话人ID和文本 transcriber.Transcribed += (s, e) => { if (e.Result.Reason == ResultReason.RecognizedSpeech) { // 从结果中获取说话人ID var speakerId = e.Result.Properties.GetProperty(PropertyId.ConversationTranscription_SpeakerId); recognizedTextBox.Invoke((MethodInvoker)delegate { recognizedTextBox.AppendText($"Speaker ID: {speakerId}, Text: {e.Result.Text}{Environment.NewLine}"); }); } }; // 订阅会话开始/结束事件(可选) transcriber.SessionStarted += (s, e) => { recognizedTextBox.Invoke((MethodInvoker)delegate { recognizedTextBox.AppendText("会话开始转录..."); }); }; transcriber.SessionStopped += (s, e) => { recognizedTextBox.Invoke((MethodInvoker)delegate { recognizedTextBox.AppendText("会话转录结束"); }); }; // 开始转录 await transcriber.StartTranscribingAsync(); // 等待转录完成(根据音频长度调整,或者监听SessionStopped事件) await Task.Delay(TimeSpan.FromSeconds(100)); // 停止转录 await transcriber.StopTranscribingAsync(); } } catch (Exception ex) { MessageBox.Show($"发生错误: {ex.Message}", "错误", MessageBoxButtons.OK, MessageBoxIcon.Error); } }
3. 额外说明
- 对话转录支持离线处理预录音频,也支持实时流音频。
- 如果需要给说话人ID关联具体名称,可以在识别完成后,通过说话人识别API将ID与已知说话人匹配;如果是未知说话人,只能暂时用返回的ID区分,后续可以手动标注名称。
- 确保音频质量良好,说话人语音清晰,否则会影响说话人区分的准确性。
内容的提问来源于stack exchange,提问作者vampire
相关产品推荐
相关产品推荐

