You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过.NET(C#)在Microsoft Bot Framework模拟器中获取客户端语音流

刚好我之前基于Microsoft Bot Framework做过实时语音转文本的场景,用.NET(C#)实现你的需求可以分客户端(测试端)和服务器端两部分来落地,下面给你一步步拆解细节:


一、先搞定依赖包

不管是客户端还是服务器端,先把这些NuGet包装上:

  • Microsoft.Bot.Builder:Bot Framework核心库
  • Microsoft.Bot.Streaming:处理实时流传输
  • Azure.AI.Speech:微软官方的语音处理SDK,用来做语音捕获、说话检测和转文本
  • Microsoft.Bot.Builder.Integration.AspNet.Core:服务器端Web应用的Bot集成包

二、客户端(自定义测试端,模拟Bot模拟器的语音输入)

官方Bot模拟器的代码我们没法修改,所以如果要在客户端侧主动捕获语音流、检测用户说话状态,得自己写个简单的测试客户端。这里用Azure Speech SDK来搞定语音相关的逻辑:

2.1 捕获语音流+说话状态检测

using Azure.AI.Speech;
using Azure.AI.Speech.Audio;
using Microsoft.Bot.Streaming;
using Microsoft.Bot.Streaming.Transport.Tcp;

// 初始化Speech配置——记得换成你自己的Azure Speech密钥和区域
var speechConfig = SpeechConfig.FromSubscription("你的Azure Speech密钥", "你的区域");
speechConfig.SpeechRecognitionLanguage = "zh-CN"; // 支持多语言,按需调整

// 绑定麦克风作为音频输入源
using var audioConfig = AudioConfig.FromDefaultMicrophoneInput();
using var speechRecognizer = new SpeechRecognizer(speechConfig, audioConfig);

// 监听用户开始/停止说话的事件
speechRecognizer.SessionStarted += (sender, args) =>
{
    Console.WriteLine("👉 用户开始说话了");
};

speechRecognizer.SessionStopped += (sender, args) =>
{
    Console.WriteLine("👈 用户停止说话了");
};

// 连接到机器人服务器的Streaming端点
var botTcpEndpoint = new Uri("tcp://localhost:3978/v3/conversations/test-conversation/stream");
var streamingClient = new TcpClientTransport(botTcpEndpoint);
await streamingClient.ConnectAsync();

// 创建发送语音流的请求
var voiceRequest = new StreamingRequest()
{
    Path = "/api/voice-stream",
    Verb = "POST"
};

// 实时推送麦克风音频到服务器
using var pushStream = AudioInputStream.CreatePushStream();
using var streamAudioConfig = AudioConfig.FromStreamInput(pushStream);
using var streamingSpeechRecognizer = new SpeechRecognizer(speechConfig, streamAudioConfig);

// 把麦克风捕获的音频数据实时推送到机器人服务器
pushStream.DataAvailable += async (sender, args) =>
{
    await voiceRequest.SendStreamAsync(args.Data);
};

// 启动语音识别(同时开始推送流)
await streamingSpeechRecognizer.StartContinuousRecognitionAsync();

Console.WriteLine("🎤 可以开始说话了,按任意键退出...");
Console.ReadKey();

// 清理资源
await streamingSpeechRecognizer.StopContinuousRecognitionAsync();
await streamingClient.DisconnectAsync();

2.2 逻辑说明

这里用Azure Speech SDK的AudioInputStream捕获麦克风的实时音频,通过SessionStarted/SessionStopped事件精准判断用户的说话状态;再借助Bot Framework的TcpClientTransport把音频流发送到机器人服务器的Streaming端点。


三、服务器端(机器人Web应用)接收语音流并转文本

服务器端需要先配置Bot的Streaming支持,再编写逻辑接收语音流并调用语音API转文本:

3.1 配置Bot的Streaming端点(以.NET 6+为例)

在Program.cs里添加Bot和Streaming的配置:

var builder = WebApplication.CreateBuilder(args);

// 添加Bot服务
builder.Services.AddHttpClient()
                .AddControllers()
                .AddNewtonsoftJson();

builder.Services.AddBot<VoiceProcessingBot>(options =>
{
    var appId = builder.Configuration["MicrosoftAppId"];
    var appPassword = builder.Configuration["MicrosoftAppPassword"];
    if (!string.IsNullOrEmpty(appPassword))
    {
        options.CredentialProvider = new SimpleCredentialProvider(appId, appPassword);
    }
});

// 启用Streaming支持
builder.Services.AddBotStreaming(options =>
{
    options.Server = new WebSocketServer(options);
});

var app = builder.Build();

app.UseRouting();
app.UseEndpoints(endpoints =>
{
    endpoints.MapControllers();
    endpoints.MapBotStreaming(); // 映射Streaming端点
});

app.Run();

3.2 实现Bot的语音处理逻辑

创建VoiceProcessingBot.cs,继承ActivityHandler,处理收到的语音流:

using Microsoft.Bot.Builder;
using Microsoft.Bot.Schema;
using Azure.AI.Speech;
using Azure.AI.Speech.Audio;

public class VoiceProcessingBot : ActivityHandler
{
    private readonly SpeechConfig _speechConfig;

    // 通过依赖注入获取配置
    public VoiceProcessingBot(IConfiguration configuration)
    {
        _speechConfig = SpeechConfig.FromSubscription(
            configuration["AzureSpeechKey"], 
            configuration["AzureSpeechRegion"]
        );
        _speechConfig.SpeechRecognitionLanguage = "zh-CN";
    }

    protected override async Task OnMessageActivityAsync(ITurnContext<IMessageActivity> turnContext, CancellationToken cancellationToken)
    {
        // 检查消息是否包含语音附件
        if (turnContext.Activity.Attachments?.Any(a => a.ContentType.StartsWith("audio/")) == true)
        {
            var voiceAttachment = turnContext.Activity.Attachments.First();
            // 从附件中获取语音流
            using var audioStream = await turnContext.TurnState
                .Get<IConnectorClient>()
                .Attachments
                .GetAttachmentAsync(voiceAttachment.ContentUrl);

            // 调用Azure Speech SDK把语音流转为文本
            using var audioConfig = AudioConfig.FromStreamInput(audioStream);
            using var speechRecognizer = new SpeechRecognizer(_speechConfig, audioConfig);

            var recognitionResult = await speechRecognizer.RecognizeOnceAsync();
            if (recognitionResult.Reason == ResultReason.RecognizedSpeech)
            {
                await turnContext.SendActivityAsync(
                    MessageFactory.Text($"✅ 识别结果:{recognitionResult.Text}"), 
                    cancellationToken
                );
            }
            else
            {
                await turnContext.SendActivityAsync(
                    MessageFactory.Text("❌ 没听清你说什么,请再试一次~"), 
                    cancellationToken
                );
            }
        }
        else
        {
            await turnContext.SendActivityAsync(
                MessageFactory.Text("请发送语音消息哦"), 
                cancellationToken
            );
        }
    }
}

3.3 实时边说边转的进阶玩法

如果需要实时返回识别结果(用户说话时就同步输出文本),可以用连续识别功能:

protected override async Task OnMessageActivityAsync(ITurnContext<IMessageActivity> turnContext, CancellationToken cancellationToken)
{
    if (turnContext.Activity.Attachments?.Any(a => a.ContentType.StartsWith("audio/")) == true)
    {
        var voiceAttachment = turnContext.Activity.Attachments.First();
        using var audioStream = await turnContext.TurnState
            .Get<IConnectorClient>()
            .Attachments
            .GetAttachmentAsync(voiceAttachment.ContentUrl);

        using var pushStream = AudioInputStream.CreatePushStream();
        using var audioConfig = AudioConfig.FromStreamInput(pushStream);
        using var speechRecognizer = new SpeechRecognizer(_speechConfig, audioConfig);

        // 实时返回中间识别结果
        speechRecognizer.Recognizing += async (sender, args) =>
        {
            if (args.Result.Reason == ResultReason.RecognizingSpeech)
            {
                await turnContext.SendActivityAsync(
                    MessageFactory.Text($"🔄 实时识别:{args.Result.Text}"), 
                    cancellationToken
                );
            }
        };

        // 返回最终识别结果
        speechRecognizer.Recognized += async (sender, args) =>
        {
            if (args.Result.Reason == ResultReason.RecognizedSpeech)
            {
                await turnContext.SendActivityAsync(
                    MessageFactory.Text($"✅ 最终结果:{args.Result.Text}"), 
                    cancellationToken
                );
            }
        };

        // 把音频流推送到Speech SDK
        byte[] buffer = new byte[1024];
        int bytesRead;
        while ((bytesRead = await audioStream.ReadAsync(buffer, 0, buffer.Length)) > 0)
        {
            pushStream.Write(buffer, 0, bytesRead);
        }

        await speechRecognizer.StartContinuousRecognitionAsync();
        await Task.Delay(5000); // 等待识别完成
        await speechRecognizer.StopContinuousRecognitionAsync();
    }
}

四、用官方Bot模拟器测试的简化方案

如果不想自己开发客户端,直接用官方Bot模拟器也能测试:

  1. 启动机器人Web应用,打开Bot模拟器,连接到http://localhost:3978/api/messages
  2. 点击模拟器的语音按钮,说话后发送
  3. 服务器端的OnMessageActivityAsync会自动收到带语音附件的消息,按上面的代码处理即可

这种情况下,说话状态的检测可以在服务器端通过Speech SDK的SessionStarted/SessionStopped事件实现。


几个注意点

  • 一定要确保Azure Speech服务的密钥和区域配置正确,不然调用会失败
  • 机器人服务必须启用Streaming支持,否则没法接收实时语音流
  • 生产环境要加异常处理,比如网络中断、音频格式不兼容的情况

内容的提问来源于stack exchange,提问作者Ariful Islam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:46:10