You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用C#与IBM Watson SDK实现无本地存储的实时音频流转文字

实现IBM Watson Speech-to-Text的实时流式语音识别(无本地文件存储)

没问题,我来帮你把现有代码改成实时流式传输的模式,不用再先存本地文件啦!IBM Watson的Speech-to-Text SDK确实支持流式识别,只是你之前用的Recognize是批量处理接口,我们需要用到它的会话式流式API来实现实时传输。

核心思路

我们会直接把NAudio录制的音频数据,通过Watson SDK提供的RecognizeSession实时推送到服务端,同时订阅识别结果的事件,实时获取转写内容,全程不需要写入本地文件。

修改后的完整代码

using System;
using System.IO;
using System.Threading.Tasks;
using IBM.Cloud.SDK.Core.Authentication.Iam;
using IBM.Watson.SpeechToText.v1;
using IBM.Watson.SpeechToText.v1.Model;
using NAudio.Wave;

namespace watsonConsole
{
    class Program
    {
        // 请替换成你的Watson凭证和服务地址
        private readonly string sttApiKey = "YOUR_STT_API_KEY";
        private readonly string sttUrl = "YOUR_STT_SERVICE_URL";
        private readonly string versionDate = "2023-05-29"; // 使用最新的版本日期

        private WaveInEvent _waveIn;
        private RecognizeSession _recognizeSession;
        private bool _isRecording = false;

        static void Main(string[] args)
        {
            Program pr = new Program();
            pr.StartStreamingRecognition().Wait();

            Console.WriteLine("Press any key to exit...");
            Console.ReadKey();
        }

        public async Task StartStreamingRecognition()
        {
            // 初始化Watson Speech-to-Text认证和服务
            var authenticator = new IamAuthenticator(apikey: sttApiKey);
            var speechToTextService = new SpeechToTextService(versionDate, authenticator);
            speechToTextService.SetServiceUrl(sttUrl);

            // 设置音频格式:推荐16kHz、16位、单声道(和Watson服务兼容)
            var waveFormat = new WaveFormat(16000, 16, 1);

            // 创建识别会话,配置流式参数
            _recognizeSession = speechToTextService.CreateRecognizeSession(
                new RecognizeOptions
                {
                    ContentType = $"audio/wav; rate={waveFormat.SampleRate}",
                    SmartFormatting = true,
                    WordAlternativesThreshold = 0.9f
                });

            // 订阅实时识别结果事件
            _recognizeSession.OnRecognize += (sender, result) =>
            {
                if (result != null && result.Results != null)
                {
                    foreach (var resultItem in result.Results)
                    {
                        if (resultItem.Final)
                        {
                            // 最终识别结果
                            Console.WriteLine($"Final Result: {resultItem.Alternatives[0].Transcript}");
                        }
                        else
                        {
                            // 中间临时结果(实时更新)
                            Console.WriteLine($"Partial Result: {resultItem.Alternatives[0].Transcript}");
                        }
                    }
                }
            };

            // 初始化NAudio音频录制
            _waveIn = new WaveInEvent
            {
                BufferMilliseconds = 100, // 每100ms推送一次音频数据,平衡实时性和性能
                DeviceNumber = 0,
                WaveFormat = waveFormat
            };

            // 音频数据可用时,直接写入Watson会话的流
            _waveIn.DataAvailable += async (sender, e) =>
            {
                if (_isRecording && _recognizeSession.AudioStream != null)
                {
                    await _recognizeSession.AudioStream.WriteAsync(e.Buffer, 0, e.BytesRecorded);
                    await _recognizeSession.AudioStream.FlushAsync();
                }
            };

            // 开始录制和识别
            Console.WriteLine("Starting streaming recording... Press Enter to stop.");
            _isRecording = true;
            _waveIn.StartRecording();
            await _recognizeSession.StartListeningAsync();

            // 等待用户输入停止
            Console.ReadLine();

            // 停止录制和会话
            _isRecording = false;
            _waveIn.StopRecording();
            await _recognizeSession.StopListeningAsync();

            // 释放资源
            _waveIn.Dispose();
            _recognizeSession.Dispose();
        }
    }
}

关键部分说明

  1. 认证与服务初始化:用IamAuthenticator替代你之前的BearerTokenAuthenticator,更适合长期使用(会自动处理token刷新)。
  2. RecognizeSession:这是SDK实现流式识别的核心对象,它内部已经封装了WebSocket连接逻辑,不用你手动处理WebSocket细节。
  3. 音频格式匹配:Watson服务对音频格式有严格要求,这里用16000Hz、16位、单声道是最稳定的配置,要确保NAudio的录制格式和ContentType参数完全一致。
  4. 实时数据推送:在NAudio的DataAvailable事件中,直接把音频缓冲区的数据写入_recognizeSession.AudioStream,实现无文件的实时传输。
  5. 结果订阅:通过OnRecognize事件可以获取中间的部分结果(实时更新的转写内容)和最终的确定结果(完成一段语音后的最终转写),满足实时转写的需求。

注意事项

  • 请确保你的Watson Speech-to-Text服务已开通,并且sttApiKey和sttUrl凭证正确。
  • 音频格式必须和服务端要求一致,否则会出现识别失败或结果不准确的情况。
  • 停止录制后一定要调用StopListeningAsync并释放资源,避免WebSocket连接泄漏。

内容的提问来源于stack exchange,提问作者Nithin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 00:47:36