You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Azure Cognitive Services生成区分说话人的JSON格式对话?

Azure认知服务说话人区分(C#实现)

要实现Azure认知服务的说话人区分并输出结构化JSON/按说话人拆分文本,你需要启用**说话人区分(Speaker Diarization)**功能,以下是完整的C#实现步骤及代码:

1. 安装依赖

首先安装NuGet包:

Install-Package Microsoft.CognitiveServices.Speech

2. 核心实现代码

using Microsoft.CognitiveServices.Speech;
using Microsoft.CognitiveServices.Speech.Audio;
using System.Text.Json;
using System.Collections.Generic;
using System.Threading.Tasks;
using System.IO;

class SpeakerDiarizationDemo
{
    static async Task Main(string[] args)
    {
        // 替换为你的Azure语音服务订阅密钥和区域
        string subscriptionKey = "YOUR_SUBSCRIPTION_KEY";
        string region = "YOUR_REGION";
        // 替换为你的音频文件路径(或从存储账户获取的流)
        string audioFilePath = "path/to/your/audio.wav";

        // 配置语音服务,启用说话人区分
        var speechConfig = SpeechConfig.FromSubscription(subscriptionKey, region);
        speechConfig.SetProperty("SpeechServiceConnection_DiarizationEnabled", "true");
        // 指定对话中的说话人数量(可选,能提升识别精度)
        speechConfig.SetProperty("SpeechServiceConnection_NumberOfSpeakers", "2");

        // 加载音频文件(如果是从Azure存储读取,可用AudioConfig.FromStreamInput()传入Blob流)
        using var audioConfig = AudioConfig.FromWavFileInput(audioFilePath);
        using var recognizer = new SpeechRecognizer(speechConfig, audioConfig);

        // 存储每个说话人的语句
        var speakerUtterances = new List<(string SpeakerId, string Text)>();

        // 监听识别结果事件
        recognizer.Recognized += (sender, e) =>
        {
            if (e.Result.Reason == ResultReason.RecognizedSpeech)
            {
                // 提取说话人ID
                string speakerId = e.Result.Properties.GetProperty(PropertyId.SpeechServiceConnection_DiarizationSpeakerId);
                speakerUtterances.Add((speakerId, e.Result.Text));
            }
        };

        // 启动连续识别(适合长音频)
        await recognizer.StartContinuousRecognitionAsync();
        // 等待识别完成(可根据音频时长调整,或监听SessionStopped事件)
        await Task.Delay(TimeSpan.FromSeconds(30));
        await recognizer.StopContinuousRecognitionAsync();

        // --- 按对话顺序合并同一说话人的连续语句 ---
        var mergedDialog = new List<(string SpeakerId, string FullText)>();
        foreach (var utterance in speakerUtterances)
        {
            if (mergedDialog.Count == 0 || mergedDialog.Last().SpeakerId != utterance.SpeakerId)
            {
                mergedDialog.Add((utterance.SpeakerId, utterance.Text));
            }
            else
            {
                var lastItem = mergedDialog.Last();
                mergedDialog.RemoveAt(mergedDialog.Count - 1);
                mergedDialog.Add((lastItem.SpeakerId, $"{lastItem.FullText} {utterance.Text}"));
            }
        }

        // 输出你期望的格式
        Console.WriteLine("按说话人拆分的对话:");
        foreach (var item in mergedDialog)
        {
            Console.WriteLine(item.FullText);
        }

        // --- 生成JSON格式输出 ---
        var jsonResult = JsonSerializer.Serialize(mergedDialog, new JsonSerializerOptions { WriteIndented = true });
        Console.WriteLine("\nJSON格式对话结果:");
        Console.WriteLine(jsonResult);

        // 保存JSON到文件
        File.WriteAllText("对话结果.json", jsonResult);
    }
}

关键说明

  • 说话人区分配置:必须设置SpeechServiceConnection_DiarizationEnabled为true,指定NumberOfSpeakers能提升识别准确率
  • 说话人ID提取:通过PropertyId.SpeechServiceConnection_DiarizationSpeakerId从识别结果中获取唯一说话人标识
  • 音频输入:支持本地文件、内存流(适配Azure存储Blob的读取场景)
  • 连续识别模式:适合处理较长的通话录音,避免单次识别长度限制

内容的提问来源于stack exchange,提问作者pdx_boats

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 22:50:32