如何通过.NET(C#)在Microsoft Bot Framework模拟器中获取客户端语音流
刚好我之前基于Microsoft Bot Framework做过实时语音转文本的场景,用.NET(C#)实现你的需求可以分客户端(测试端)和服务器端两部分来落地,下面给你一步步拆解细节:
一、先搞定依赖包
不管是客户端还是服务器端,先把这些NuGet包装上:
Microsoft.Bot.Builder:Bot Framework核心库Microsoft.Bot.Streaming:处理实时流传输Azure.AI.Speech:微软官方的语音处理SDK,用来做语音捕获、说话检测和转文本Microsoft.Bot.Builder.Integration.AspNet.Core:服务器端Web应用的Bot集成包
二、客户端(自定义测试端,模拟Bot模拟器的语音输入)
官方Bot模拟器的代码我们没法修改,所以如果要在客户端侧主动捕获语音流、检测用户说话状态,得自己写个简单的测试客户端。这里用Azure Speech SDK来搞定语音相关的逻辑:
2.1 捕获语音流+说话状态检测
using Azure.AI.Speech; using Azure.AI.Speech.Audio; using Microsoft.Bot.Streaming; using Microsoft.Bot.Streaming.Transport.Tcp; // 初始化Speech配置——记得换成你自己的Azure Speech密钥和区域 var speechConfig = SpeechConfig.FromSubscription("你的Azure Speech密钥", "你的区域"); speechConfig.SpeechRecognitionLanguage = "zh-CN"; // 支持多语言,按需调整 // 绑定麦克风作为音频输入源 using var audioConfig = AudioConfig.FromDefaultMicrophoneInput(); using var speechRecognizer = new SpeechRecognizer(speechConfig, audioConfig); // 监听用户开始/停止说话的事件 speechRecognizer.SessionStarted += (sender, args) => { Console.WriteLine("👉 用户开始说话了"); }; speechRecognizer.SessionStopped += (sender, args) => { Console.WriteLine("👈 用户停止说话了"); }; // 连接到机器人服务器的Streaming端点 var botTcpEndpoint = new Uri("tcp://localhost:3978/v3/conversations/test-conversation/stream"); var streamingClient = new TcpClientTransport(botTcpEndpoint); await streamingClient.ConnectAsync(); // 创建发送语音流的请求 var voiceRequest = new StreamingRequest() { Path = "/api/voice-stream", Verb = "POST" }; // 实时推送麦克风音频到服务器 using var pushStream = AudioInputStream.CreatePushStream(); using var streamAudioConfig = AudioConfig.FromStreamInput(pushStream); using var streamingSpeechRecognizer = new SpeechRecognizer(speechConfig, streamAudioConfig); // 把麦克风捕获的音频数据实时推送到机器人服务器 pushStream.DataAvailable += async (sender, args) => { await voiceRequest.SendStreamAsync(args.Data); }; // 启动语音识别(同时开始推送流) await streamingSpeechRecognizer.StartContinuousRecognitionAsync(); Console.WriteLine("🎤 可以开始说话了,按任意键退出..."); Console.ReadKey(); // 清理资源 await streamingSpeechRecognizer.StopContinuousRecognitionAsync(); await streamingClient.DisconnectAsync();
2.2 逻辑说明
这里用Azure Speech SDK的AudioInputStream捕获麦克风的实时音频,通过SessionStarted/SessionStopped事件精准判断用户的说话状态;再借助Bot Framework的TcpClientTransport把音频流发送到机器人服务器的Streaming端点。
三、服务器端(机器人Web应用)接收语音流并转文本
服务器端需要先配置Bot的Streaming支持,再编写逻辑接收语音流并调用语音API转文本:
3.1 配置Bot的Streaming端点(以.NET 6+为例)
在Program.cs里添加Bot和Streaming的配置:
var builder = WebApplication.CreateBuilder(args); // 添加Bot服务 builder.Services.AddHttpClient() .AddControllers() .AddNewtonsoftJson(); builder.Services.AddBot<VoiceProcessingBot>(options => { var appId = builder.Configuration["MicrosoftAppId"]; var appPassword = builder.Configuration["MicrosoftAppPassword"]; if (!string.IsNullOrEmpty(appPassword)) { options.CredentialProvider = new SimpleCredentialProvider(appId, appPassword); } }); // 启用Streaming支持 builder.Services.AddBotStreaming(options => { options.Server = new WebSocketServer(options); }); var app = builder.Build(); app.UseRouting(); app.UseEndpoints(endpoints => { endpoints.MapControllers(); endpoints.MapBotStreaming(); // 映射Streaming端点 }); app.Run();
3.2 实现Bot的语音处理逻辑
创建VoiceProcessingBot.cs,继承ActivityHandler,处理收到的语音流:
using Microsoft.Bot.Builder; using Microsoft.Bot.Schema; using Azure.AI.Speech; using Azure.AI.Speech.Audio; public class VoiceProcessingBot : ActivityHandler { private readonly SpeechConfig _speechConfig; // 通过依赖注入获取配置 public VoiceProcessingBot(IConfiguration configuration) { _speechConfig = SpeechConfig.FromSubscription( configuration["AzureSpeechKey"], configuration["AzureSpeechRegion"] ); _speechConfig.SpeechRecognitionLanguage = "zh-CN"; } protected override async Task OnMessageActivityAsync(ITurnContext<IMessageActivity> turnContext, CancellationToken cancellationToken) { // 检查消息是否包含语音附件 if (turnContext.Activity.Attachments?.Any(a => a.ContentType.StartsWith("audio/")) == true) { var voiceAttachment = turnContext.Activity.Attachments.First(); // 从附件中获取语音流 using var audioStream = await turnContext.TurnState .Get<IConnectorClient>() .Attachments .GetAttachmentAsync(voiceAttachment.ContentUrl); // 调用Azure Speech SDK把语音流转为文本 using var audioConfig = AudioConfig.FromStreamInput(audioStream); using var speechRecognizer = new SpeechRecognizer(_speechConfig, audioConfig); var recognitionResult = await speechRecognizer.RecognizeOnceAsync(); if (recognitionResult.Reason == ResultReason.RecognizedSpeech) { await turnContext.SendActivityAsync( MessageFactory.Text($"✅ 识别结果:{recognitionResult.Text}"), cancellationToken ); } else { await turnContext.SendActivityAsync( MessageFactory.Text("❌ 没听清你说什么,请再试一次~"), cancellationToken ); } } else { await turnContext.SendActivityAsync( MessageFactory.Text("请发送语音消息哦"), cancellationToken ); } } }
3.3 实时边说边转的进阶玩法
如果需要实时返回识别结果(用户说话时就同步输出文本),可以用连续识别功能:
protected override async Task OnMessageActivityAsync(ITurnContext<IMessageActivity> turnContext, CancellationToken cancellationToken) { if (turnContext.Activity.Attachments?.Any(a => a.ContentType.StartsWith("audio/")) == true) { var voiceAttachment = turnContext.Activity.Attachments.First(); using var audioStream = await turnContext.TurnState .Get<IConnectorClient>() .Attachments .GetAttachmentAsync(voiceAttachment.ContentUrl); using var pushStream = AudioInputStream.CreatePushStream(); using var audioConfig = AudioConfig.FromStreamInput(pushStream); using var speechRecognizer = new SpeechRecognizer(_speechConfig, audioConfig); // 实时返回中间识别结果 speechRecognizer.Recognizing += async (sender, args) => { if (args.Result.Reason == ResultReason.RecognizingSpeech) { await turnContext.SendActivityAsync( MessageFactory.Text($"🔄 实时识别:{args.Result.Text}"), cancellationToken ); } }; // 返回最终识别结果 speechRecognizer.Recognized += async (sender, args) => { if (args.Result.Reason == ResultReason.RecognizedSpeech) { await turnContext.SendActivityAsync( MessageFactory.Text($"✅ 最终结果:{args.Result.Text}"), cancellationToken ); } }; // 把音频流推送到Speech SDK byte[] buffer = new byte[1024]; int bytesRead; while ((bytesRead = await audioStream.ReadAsync(buffer, 0, buffer.Length)) > 0) { pushStream.Write(buffer, 0, bytesRead); } await speechRecognizer.StartContinuousRecognitionAsync(); await Task.Delay(5000); // 等待识别完成 await speechRecognizer.StopContinuousRecognitionAsync(); } }
四、用官方Bot模拟器测试的简化方案
如果不想自己开发客户端,直接用官方Bot模拟器也能测试:
- 启动机器人Web应用,打开Bot模拟器,连接到
http://localhost:3978/api/messages - 点击模拟器的语音按钮,说话后发送
- 服务器端的
OnMessageActivityAsync会自动收到带语音附件的消息,按上面的代码处理即可
这种情况下,说话状态的检测可以在服务器端通过Speech SDK的SessionStarted/SessionStopped事件实现。
几个注意点
- 一定要确保Azure Speech服务的密钥和区域配置正确,不然调用会失败
- 机器人服务必须启用Streaming支持,否则没法接收实时语音流
- 生产环境要加异常处理,比如网络中断、音频格式不兼容的情况
内容的提问来源于stack exchange,提问作者Ariful Islam
相关产品推荐
相关产品推荐

