如何将Azure文本转语音实时音频流发送至Ozeki VoIP SIP SDK?
问题:Azure TTS音频流接入Ozeki VoIP通话
我正在开展一个项目,需使用Azure Text-to-Speech服务将文本转换为语音,再通过Ozeki VoIP SIP SDK将该语音音频实时流式传输至VoIP通话中。
已成功通过Azure生成语音音频并获取到字节数组,但无法将这些音频数据以可流式传输至VoIP通话的方式发送给Ozeki。需要将字节数组转换为Ozeki兼容的格式,并实现音频数据的实时流式传输。
尝试将Azure TTS返回的字节数组转换为MemoryStream,再通过NAudio库将其转为WaveStream,意图在通话中播放该WaveStream,但不确定如何正确将WaveStream与通话连接,也不确定这种实现实时音频流的方式是否正确。
尝试的代码
TextToSpeech类
using System; using Microsoft.CognitiveServices.Speech.Audio; using Microsoft.CognitiveServices.Speech; using System.IO; using System.Threading.Tasks; using NAudio.Wave; namespace Adion.Media { public class TextToSpeech { public async Task Speak(string text) { // create speech config var config = SpeechConfig.FromSubscription(az_key, az_reg); // create ssml var ssml = $@"<speak version='1.0' xml:lang='fr-FR' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:emo='http://www.w3.org/2009/10/emotionml' xmlns:mstts='http://www.w3.org/2001/mstts'><voice name='{az_voice}'><s /><mstts:express-as style='cheerful'>{text}</mstts:express-as><s /></voice></speak> "; // Creates an audio out stream. using (var stream = AudioOutputStream.CreatePullStream()) { // Creates a speech synthesizer using audio stream output. using (var streamConfig = AudioConfig.FromStreamOutput(stream)) using (var synthesizer = new SpeechSynthesizer(config, streamConfig)) { while (true) { // Receives a text from console input and synthesize it to pull audio output stream. if (string.IsNullOrEmpty(text)) { break; } using (var result = await synthesizer.SpeakTextAsync(text)) { if (result.Reason == ResultReason.SynthesizingAudioCompleted) { Console.WriteLine($"Speech synthesized for text [{text}], and the audio was written to output stream."); text = null; } else if (result.Reason == ResultReason.Canceled) { var cancellation = SpeechSynthesisCancellationDetails.FromResult(result); Console.WriteLine($"CANCELED: Reason={cancellation.Reason}"); if (cancellation.Reason == CancellationReason.Error) { Console.WriteLine($"CANCELED: ErrorCode={cancellation.ErrorCode}"); Console.WriteLine($"CANCELED: ErrorDetails=[{cancellation.ErrorDetails}]"); Console.WriteLine($"CANCELED: Did you update the subscription info?"); } } } } } // Reads(pulls) data from the stream byte[] buffer = new byte[32000]; uint filledSize = 0; uint totalSize = 0; MemoryStream memoryStream = new MemoryStream(); while ((filledSize = stream.Read(buffer)) > 0) { Console.WriteLine($"{filledSize} bytes received."); totalSize += filledSize; memoryStream.Write(buffer, 0, (int)filledSize); } Console.WriteLine($"Totally {totalSize} bytes received."); // Convert the MemoryStream to WaveStream WaveStream waveStream = new RawSourceWaveStream(memoryStream, new NAudio.Wave.WaveFormat()); } } } }
通话处理程序
using Ozeki.VoIP; using Ozeki.Media; using Adion.Tools; using Adion.Media; using TextToSpeech = Adion.Media.TextToSpeech; namespace Adion.SIP { internal class call_handler { static MediaConnector connector = new MediaConnector(); static PhoneCallAudioSender mediaSender = new PhoneCallAudioSender(); public static void incoming_call(object sender, VoIPEventArgs<IPhoneCall> e) { var call = e.Item; Log.info("Incoming call from: " + call.DialInfo.CallerID); call.CallStateChanged += on_call_state_changed; call.Answer(); } public static async void on_call_state_changed(object sender, CallStateChangedArgs e) { var call = sender as IPhoneCall; switch (e.State) { case CallState.Answered: Log.info("Call is answered"); break; case CallState.Completed: Log.info("Call is completed"); break; case CallState.InCall: Log.info("Call is in progress"); var textToSpeech = new TextToSpeech(); mediaSender.AttachToCall(call); connector.Connect(textToSpeech, mediaSender); textToSpeech.AddAndStartText("I can't understand why this texte can be hear in the voip cal !!!"); break; } } } }
解决建议
1. 对齐音频格式
Ozeki VoIP通话默认使用PCM 8kHz 16位单声道音频格式,需强制Azure TTS输出对应格式,避免格式不兼容:
var config = SpeechConfig.FromSubscription(az_key, az_reg); // 指定输出格式为Ozeki兼容的PCM config.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Raw16Khz16BitMonoPcm);
2. 实现实时流式推送
当前代码先缓存所有音频再处理,不符合实时需求。改用PushStream实时获取音频数据并推送给Ozeki:
using Ozeki.Media; using NAudio.Wave; namespace Adion.Media { public class TextToSpeech : IAudioSender { public event EventHandler<AudioDataReceivedEventArgs> AudioDataReceived; private readonly SpeechConfig _speechConfig; private readonly string _voiceName; public TextToSpeech(string azKey, string azRegion, string azVoice) { _speechConfig = SpeechConfig.FromSubscription(azKey, azRegion); _speechConfig.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Raw16Khz16BitMonoPcm); _voiceName = azVoice; } public async Task Speak(string text) { var ssml = $@"<speak version='1.0' xml:lang='fr-FR' xmlns='http://www.w3.org/2001/10/synthesis' xmlns:mstts='http://www.w3.org/2001/mstts'><voice name='{_voiceName}'><s /><mstts:express-as style='cheerful'>{text}</mstts:express-as><s /></voice></speak>"; using var stream = AudioOutputStream.CreatePushStream(); using var streamConfig = AudioConfig.FromStreamOutput(stream); using var synthesizer = new SpeechSynthesizer(_speechConfig, streamConfig); // 实时推送音频数据 stream.DataAvailable += (sender, args) => { var audioData = new AudioData(args.Data, args.Length); AudioDataReceived?.Invoke(this, new AudioDataReceivedEventArgs(audioData)); }; var result = await synthesizer.SpeakSsmlAsync(ssml); if (result.Reason == ResultReason.Canceled) { var cancellation = SpeechSynthesisCancellationDetails.FromResult(result); Console.WriteLine($"合成失败: {cancellation.Reason}, 详情: {cancellation.ErrorDetails}"); } } // 实现IAudioSender接口方法 public void Start() {} public void Stop() {} public void Dispose() {} } }
3. 正确连接Ozeki媒体链路
修改通话处理逻辑,确保TTS实例与Ozeki媒体发送器正确连接:
namespace Adion.SIP { internal class call_handler { static MediaConnector connector = new MediaConnector(); static PhoneCallAudioSender mediaSender = new PhoneCallAudioSender(); public static void incoming_call(object sender, VoIPEventArgs<IPhoneCall> e) { var call = e.Item; Log.info("来电号码: " + call.DialInfo.CallerID); call.CallStateChanged += on_call_state_changed; call.Answer(); } public static async void on_call_state_changed(object sender, CallStateChangedArgs e) { var call = sender as IPhoneCall; switch (e.State) { case CallState.Answered: Log.info("通话已接听"); break; case CallState.Completed: Log.info("通话已结束"); break; case CallState.InCall: Log.info("通话进行中"); // 初始化TTS实例 var textToSpeech = new TextToSpeech(az_key, az_reg, az_voice); mediaSender.AttachToCall(call); // 连接TTS到媒体发送器 connector.Connect(textToSpeech, mediaSender); // 开始合成并推送音频 await textToSpeech.Speak("现在你应该能在通话中听到这段语音了!"); break; } } } }
关键注意事项
- 必须保证Azure TTS输出格式与Ozeki通话格式完全一致,否则会出现无声或杂音
- 使用
PushStream替代PullStream,实现音频数据的实时获取与推送 - 实现Ozeki的
IAudioSender接口,让TTS类直接接入Ozeki媒体链路,无需额外格式转换
内容的提问来源于stack exchange,提问作者Adion
相关产品推荐
相关产品推荐

