如何用C#与IBM Watson SDK实现无本地存储的实时音频流转文字
实现IBM Watson Speech-to-Text的实时流式语音识别(无本地文件存储)
没问题,我来帮你把现有代码改成实时流式传输的模式,不用再先存本地文件啦!IBM Watson的Speech-to-Text SDK确实支持流式识别,只是你之前用的Recognize是批量处理接口,我们需要用到它的会话式流式API来实现实时传输。
核心思路
我们会直接把NAudio录制的音频数据,通过Watson SDK提供的RecognizeSession实时推送到服务端,同时订阅识别结果的事件,实时获取转写内容,全程不需要写入本地文件。
修改后的完整代码
using System; using System.IO; using System.Threading.Tasks; using IBM.Cloud.SDK.Core.Authentication.Iam; using IBM.Watson.SpeechToText.v1; using IBM.Watson.SpeechToText.v1.Model; using NAudio.Wave; namespace watsonConsole { class Program { // 请替换成你的Watson凭证和服务地址 private readonly string sttApiKey = "YOUR_STT_API_KEY"; private readonly string sttUrl = "YOUR_STT_SERVICE_URL"; private readonly string versionDate = "2023-05-29"; // 使用最新的版本日期 private WaveInEvent _waveIn; private RecognizeSession _recognizeSession; private bool _isRecording = false; static void Main(string[] args) { Program pr = new Program(); pr.StartStreamingRecognition().Wait(); Console.WriteLine("Press any key to exit..."); Console.ReadKey(); } public async Task StartStreamingRecognition() { // 初始化Watson Speech-to-Text认证和服务 var authenticator = new IamAuthenticator(apikey: sttApiKey); var speechToTextService = new SpeechToTextService(versionDate, authenticator); speechToTextService.SetServiceUrl(sttUrl); // 设置音频格式:推荐16kHz、16位、单声道(和Watson服务兼容) var waveFormat = new WaveFormat(16000, 16, 1); // 创建识别会话,配置流式参数 _recognizeSession = speechToTextService.CreateRecognizeSession( new RecognizeOptions { ContentType = $"audio/wav; rate={waveFormat.SampleRate}", SmartFormatting = true, WordAlternativesThreshold = 0.9f }); // 订阅实时识别结果事件 _recognizeSession.OnRecognize += (sender, result) => { if (result != null && result.Results != null) { foreach (var resultItem in result.Results) { if (resultItem.Final) { // 最终识别结果 Console.WriteLine($"Final Result: {resultItem.Alternatives[0].Transcript}"); } else { // 中间临时结果(实时更新) Console.WriteLine($"Partial Result: {resultItem.Alternatives[0].Transcript}"); } } } }; // 初始化NAudio音频录制 _waveIn = new WaveInEvent { BufferMilliseconds = 100, // 每100ms推送一次音频数据,平衡实时性和性能 DeviceNumber = 0, WaveFormat = waveFormat }; // 音频数据可用时,直接写入Watson会话的流 _waveIn.DataAvailable += async (sender, e) => { if (_isRecording && _recognizeSession.AudioStream != null) { await _recognizeSession.AudioStream.WriteAsync(e.Buffer, 0, e.BytesRecorded); await _recognizeSession.AudioStream.FlushAsync(); } }; // 开始录制和识别 Console.WriteLine("Starting streaming recording... Press Enter to stop."); _isRecording = true; _waveIn.StartRecording(); await _recognizeSession.StartListeningAsync(); // 等待用户输入停止 Console.ReadLine(); // 停止录制和会话 _isRecording = false; _waveIn.StopRecording(); await _recognizeSession.StopListeningAsync(); // 释放资源 _waveIn.Dispose(); _recognizeSession.Dispose(); } } }
关键部分说明
- 认证与服务初始化:用
IamAuthenticator替代你之前的BearerTokenAuthenticator,更适合长期使用(会自动处理token刷新)。 - RecognizeSession:这是SDK实现流式识别的核心对象,它内部已经封装了WebSocket连接逻辑,不用你手动处理WebSocket细节。
- 音频格式匹配:Watson服务对音频格式有严格要求,这里用
16000Hz、16位、单声道是最稳定的配置,要确保NAudio的录制格式和ContentType参数完全一致。 - 实时数据推送:在NAudio的
DataAvailable事件中,直接把音频缓冲区的数据写入_recognizeSession.AudioStream,实现无文件的实时传输。 - 结果订阅:通过
OnRecognize事件可以获取中间的部分结果(实时更新的转写内容)和最终的确定结果(完成一段语音后的最终转写),满足实时转写的需求。
注意事项
- 请确保你的Watson Speech-to-Text服务已开通,并且
sttApiKey和sttUrl凭证正确。 - 音频格式必须和服务端要求一致,否则会出现识别失败或结果不准确的情况。
- 停止录制后一定要调用
StopListeningAsync并释放资源,避免WebSocket连接泄漏。
内容的提问来源于stack exchange,提问作者Nithin
相关产品推荐
相关产品推荐

