如何用Azure Cognitive Services生成区分说话人的JSON格式对话?
Azure认知服务说话人区分(C#实现)
要实现Azure认知服务的说话人区分并输出结构化JSON/按说话人拆分文本,你需要启用**说话人区分(Speaker Diarization)**功能,以下是完整的C#实现步骤及代码:
1. 安装依赖
首先安装NuGet包:
Install-Package Microsoft.CognitiveServices.Speech
2. 核心实现代码
using Microsoft.CognitiveServices.Speech; using Microsoft.CognitiveServices.Speech.Audio; using System.Text.Json; using System.Collections.Generic; using System.Threading.Tasks; using System.IO; class SpeakerDiarizationDemo { static async Task Main(string[] args) { // 替换为你的Azure语音服务订阅密钥和区域 string subscriptionKey = "YOUR_SUBSCRIPTION_KEY"; string region = "YOUR_REGION"; // 替换为你的音频文件路径(或从存储账户获取的流) string audioFilePath = "path/to/your/audio.wav"; // 配置语音服务,启用说话人区分 var speechConfig = SpeechConfig.FromSubscription(subscriptionKey, region); speechConfig.SetProperty("SpeechServiceConnection_DiarizationEnabled", "true"); // 指定对话中的说话人数量(可选,能提升识别精度) speechConfig.SetProperty("SpeechServiceConnection_NumberOfSpeakers", "2"); // 加载音频文件(如果是从Azure存储读取,可用AudioConfig.FromStreamInput()传入Blob流) using var audioConfig = AudioConfig.FromWavFileInput(audioFilePath); using var recognizer = new SpeechRecognizer(speechConfig, audioConfig); // 存储每个说话人的语句 var speakerUtterances = new List<(string SpeakerId, string Text)>(); // 监听识别结果事件 recognizer.Recognized += (sender, e) => { if (e.Result.Reason == ResultReason.RecognizedSpeech) { // 提取说话人ID string speakerId = e.Result.Properties.GetProperty(PropertyId.SpeechServiceConnection_DiarizationSpeakerId); speakerUtterances.Add((speakerId, e.Result.Text)); } }; // 启动连续识别(适合长音频) await recognizer.StartContinuousRecognitionAsync(); // 等待识别完成(可根据音频时长调整,或监听SessionStopped事件) await Task.Delay(TimeSpan.FromSeconds(30)); await recognizer.StopContinuousRecognitionAsync(); // --- 按对话顺序合并同一说话人的连续语句 --- var mergedDialog = new List<(string SpeakerId, string FullText)>(); foreach (var utterance in speakerUtterances) { if (mergedDialog.Count == 0 || mergedDialog.Last().SpeakerId != utterance.SpeakerId) { mergedDialog.Add((utterance.SpeakerId, utterance.Text)); } else { var lastItem = mergedDialog.Last(); mergedDialog.RemoveAt(mergedDialog.Count - 1); mergedDialog.Add((lastItem.SpeakerId, $"{lastItem.FullText} {utterance.Text}")); } } // 输出你期望的格式 Console.WriteLine("按说话人拆分的对话:"); foreach (var item in mergedDialog) { Console.WriteLine(item.FullText); } // --- 生成JSON格式输出 --- var jsonResult = JsonSerializer.Serialize(mergedDialog, new JsonSerializerOptions { WriteIndented = true }); Console.WriteLine("\nJSON格式对话结果:"); Console.WriteLine(jsonResult); // 保存JSON到文件 File.WriteAllText("对话结果.json", jsonResult); } }
关键说明
- 说话人区分配置:必须设置
SpeechServiceConnection_DiarizationEnabled为true,指定NumberOfSpeakers能提升识别准确率 - 说话人ID提取:通过
PropertyId.SpeechServiceConnection_DiarizationSpeakerId从识别结果中获取唯一说话人标识 - 音频输入:支持本地文件、内存流(适配Azure存储Blob的读取场景)
- 连续识别模式:适合处理较长的通话录音,避免单次识别长度限制
内容的提问来源于stack exchange,提问作者pdx_boats
相关产品推荐
相关产品推荐

