You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Azure文本转语音实时音频流发送至Ozeki VoIP SIP SDK?

问题:Azure TTS音频流接入Ozeki VoIP通话

我正在开展一个项目,需使用Azure Text-to-Speech服务将文本转换为语音,再通过Ozeki VoIP SIP SDK将该语音音频实时流式传输至VoIP通话中。

已成功通过Azure生成语音音频并获取到字节数组,但无法将这些音频数据以可流式传输至VoIP通话的方式发送给Ozeki。需要将字节数组转换为Ozeki兼容的格式,并实现音频数据的实时流式传输。

尝试将Azure TTS返回的字节数组转换为MemoryStream,再通过NAudio库将其转为WaveStream,意图在通话中播放该WaveStream,但不确定如何正确将WaveStream与通话连接,也不确定这种实现实时音频流的方式是否正确。

尝试的代码

TextToSpeech类

using System;
using Microsoft.CognitiveServices.Speech.Audio;
using Microsoft.CognitiveServices.Speech;
using System.IO;
using System.Threading.Tasks;
using NAudio.Wave;

namespace Adion.Media
{
    public class TextToSpeech
    {
       
        public async Task Speak(string text)
        {
            // create speech config
            var config = SpeechConfig.FromSubscription(az_key, az_reg);

            // create ssml
            var ssml = $@"<speak version='1.0' xml:lang='fr-FR' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:emo='http://www.w3.org/2009/10/emotionml'  xmlns:mstts='http://www.w3.org/2001/mstts'><voice name='{az_voice}'><s /><mstts:express-as style='cheerful'>{text}</mstts:express-as><s /></voice></speak> ";

            // Creates an audio out stream.
            using (var stream = AudioOutputStream.CreatePullStream())
            {
                // Creates a speech synthesizer using audio stream output.
                using (var streamConfig = AudioConfig.FromStreamOutput(stream))
                using (var synthesizer = new SpeechSynthesizer(config, streamConfig))
                {
                    while (true)
                    {
                        // Receives a text from console input and synthesize it to pull audio output stream.
                        if (string.IsNullOrEmpty(text))
                        {
                            break;
                        }

                        using (var result = await synthesizer.SpeakTextAsync(text))
                        {
                            if (result.Reason == ResultReason.SynthesizingAudioCompleted)
                            {
                                Console.WriteLine($"Speech synthesized for text [{text}], and the audio was written to output stream.");
                                text = null;
                            }
                            else if (result.Reason == ResultReason.Canceled)
                            {
                                var cancellation = SpeechSynthesisCancellationDetails.FromResult(result);
                                Console.WriteLine($"CANCELED: Reason={cancellation.Reason}");

                                if (cancellation.Reason == CancellationReason.Error)
                                {
                                    Console.WriteLine($"CANCELED: ErrorCode={cancellation.ErrorCode}");
                                    Console.WriteLine($"CANCELED: ErrorDetails=[{cancellation.ErrorDetails}]");
                                    Console.WriteLine($"CANCELED: Did you update the subscription info?");
                                }
                            }
                        }
                    }
                }

                // Reads(pulls) data from the stream
                byte[] buffer = new byte[32000];
                uint filledSize = 0;
                uint totalSize = 0;
                MemoryStream memoryStream = new MemoryStream();
                while ((filledSize = stream.Read(buffer)) > 0)
                {
                    Console.WriteLine($"{filledSize} bytes received.");
                    totalSize += filledSize;
                    memoryStream.Write(buffer, 0, (int)filledSize);
                }

                Console.WriteLine($"Totally {totalSize} bytes received.");

                // Convert the MemoryStream to WaveStream
                WaveStream waveStream = new RawSourceWaveStream(memoryStream, new NAudio.Wave.WaveFormat());
                

            }
        }

    }
}

通话处理程序

using Ozeki.VoIP;
using Ozeki.Media;
using Adion.Tools;
using Adion.Media;
using TextToSpeech = Adion.Media.TextToSpeech;

namespace Adion.SIP
{
    internal class call_handler
    {

        static MediaConnector connector = new MediaConnector();
        static PhoneCallAudioSender mediaSender = new PhoneCallAudioSender();

        public static void incoming_call(object sender, VoIPEventArgs<IPhoneCall> e)
        {
            var call = e.Item;
            Log.info("Incoming call from: " + call.DialInfo.CallerID);

            call.CallStateChanged += on_call_state_changed;

            call.Answer();
        }

        public static async void on_call_state_changed(object sender, CallStateChangedArgs e) 
        {

            var call = sender as IPhoneCall;
            
            switch (e.State)
            {
                case CallState.Answered:
                    Log.info("Call is answered");
                    break;
                case CallState.Completed:
                    Log.info("Call is completed");
                    break;
                case CallState.InCall:
                    Log.info("Call is in progress");
                    
                    var textToSpeech = new TextToSpeech();
                    
                    mediaSender.AttachToCall(call);
                    connector.Connect(textToSpeech, mediaSender);

                    textToSpeech.AddAndStartText("I can't understand why this texte can be hear in the voip cal !!!");
                    
                    break;
            }
        }
    }
}

解决建议

1. 对齐音频格式

Ozeki VoIP通话默认使用PCM 8kHz 16位单声道音频格式,需强制Azure TTS输出对应格式,避免格式不兼容:

var config = SpeechConfig.FromSubscription(az_key, az_reg);
// 指定输出格式为Ozeki兼容的PCM
config.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Raw16Khz16BitMonoPcm);

2. 实现实时流式推送

当前代码先缓存所有音频再处理,不符合实时需求。改用PushStream实时获取音频数据并推送给Ozeki:

using Ozeki.Media;
using NAudio.Wave;

namespace Adion.Media
{
    public class TextToSpeech : IAudioSender
    {
        public event EventHandler<AudioDataReceivedEventArgs> AudioDataReceived;
        
        private readonly SpeechConfig _speechConfig;
        private readonly string _voiceName;

        public TextToSpeech(string azKey, string azRegion, string azVoice)
        {
            _speechConfig = SpeechConfig.FromSubscription(azKey, azRegion);
            _speechConfig.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Raw16Khz16BitMonoPcm);
            _voiceName = azVoice;
        }

        public async Task Speak(string text)
        {
            var ssml = $@"<speak version='1.0' xml:lang='fr-FR' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts'><voice name='{_voiceName}'><s /><mstts:express-as style='cheerful'>{text}</mstts:express-as><s /></voice></speak>";

            using var stream = AudioOutputStream.CreatePushStream();
            using var streamConfig = AudioConfig.FromStreamOutput(stream);
            using var synthesizer = new SpeechSynthesizer(_speechConfig, streamConfig);

            // 实时推送音频数据
            stream.DataAvailable += (sender, args) =>
            {
                var audioData = new AudioData(args.Data, args.Length);
                AudioDataReceived?.Invoke(this, new AudioDataReceivedEventArgs(audioData));
            };

            var result = await synthesizer.SpeakSsmlAsync(ssml);
            if (result.Reason == ResultReason.Canceled)
            {
                var cancellation = SpeechSynthesisCancellationDetails.FromResult(result);
                Console.WriteLine($"合成失败: {cancellation.Reason}, 详情: {cancellation.ErrorDetails}");
            }
        }

        // 实现IAudioSender接口方法
        public void Start() {}
        public void Stop() {}
        public void Dispose() {}
    }
}

3. 正确连接Ozeki媒体链路

修改通话处理逻辑,确保TTS实例与Ozeki媒体发送器正确连接:

namespace Adion.SIP
{
    internal class call_handler
    {
        static MediaConnector connector = new MediaConnector();
        static PhoneCallAudioSender mediaSender = new PhoneCallAudioSender();

        public static void incoming_call(object sender, VoIPEventArgs<IPhoneCall> e)
        {
            var call = e.Item;
            Log.info("来电号码: " + call.DialInfo.CallerID);

            call.CallStateChanged += on_call_state_changed;
            call.Answer();
        }

        public static async void on_call_state_changed(object sender, CallStateChangedArgs e) 
        {
            var call = sender as IPhoneCall;
            
            switch (e.State)
            {
                case CallState.Answered:
                    Log.info("通话已接听");
                    break;
                case CallState.Completed:
                    Log.info("通话已结束");
                    break;
                case CallState.InCall:
                    Log.info("通话进行中");
                    
                    // 初始化TTS实例
                    var textToSpeech = new TextToSpeech(az_key, az_reg, az_voice);
                    
                    mediaSender.AttachToCall(call);
                    // 连接TTS到媒体发送器
                    connector.Connect(textToSpeech, mediaSender);
                    
                    // 开始合成并推送音频
                    await textToSpeech.Speak("现在你应该能在通话中听到这段语音了!");
                    
                    break;
            }
        }
    }
}

关键注意事项

  • 必须保证Azure TTS输出格式与Ozeki通话格式完全一致,否则会出现无声或杂音
  • 使用PushStream替代PullStream,实现音频数据的实时获取与推送
  • 实现Ozeki的IAudioSender接口,让TTS类直接接入Ozeki媒体链路,无需额外格式转换

内容的提问来源于stack exchange,提问作者Adion

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 10:07:09