Azure Text to Speech调用Riff格式时如何静默获取音频数据?
问题
使用Azure认知服务的文本转语音服务,通过SSML构建请求并调用SpeakSsmlAsync函数时,选择Audio24Khz160KBitRateMonoMp3输出格式会直接返回语音数据;但选择Riff24Khz16BitMonoPcm格式时,函数会先通过扬声器播放语音,再返回数据。需要实现静默调用该Riff格式,仅获取语音数据不播放。
2024年8月3日更新,相关代码如下:
相关代码
// 初始化SpeechConfig SpeechConfig speechConfig = SpeechConfig.FromSubscription(SubscriptionKey, SubscriptionRegion); speechConfig.OutputFormat = OutputFormat.Detailed; speechConfig.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Riff24Khz16BitMonoPcm); speechConfig.SpeechSynthesisVoiceName = "de-DE-KatjaNeural"; speechConfig.SetProperty(PropertyId.Speech_LogFilename, LogServices.SpeechLogFilepath); // 创建SpeechSynthesizer并调用服务 using (SpeechSynthesizer speechSynthesizer = new SpeechSynthesizer(speechConfig)) { // 生成SSML文本 string strSsml = _textToSSML(Language, speechConfig.SpeechSynthesisVoiceName, strText); // 调用文本转语音接口 SpeechSynthesisResult speechSynthesisResult = await speechSynthesizer.SpeakSsmlAsync(strSsml); // 处理返回结果 if (speechSynthesisResult.Reason == ResultReason.SynthesizingAudioCompleted) { // 将WAV数据处理并保存为MP3 WaveFile waveFile = new WaveFile(speechSynthesisResult.AudioData); _processAndSave(waveFile, audioFilepaths); } else { // 记录错误日志 new LogServices().AddTextToSpeechError(speechSynthesisResult, strText); } }
解决方案
问题根源在于SpeechSynthesizer的默认行为:仅传入SpeechConfig实例时,SDK会自动调用系统默认音频输出设备播放合成语音。要实现静默获取数据,只需在创建SpeechSynthesizer时传入null作为AudioConfig参数,明确禁用音频播放功能。
修改后的核心代码如下:
// 创建SpeechSynthesizer时传入null禁用音频播放 using (SpeechSynthesizer speechSynthesizer = new SpeechSynthesizer(speechConfig, null)) { // 后续逻辑保持不变 string strSsml = _textToSSML(Language, speechConfig.SpeechSynthesisVoiceName, strText); SpeechSynthesisResult speechSynthesisResult = await speechSynthesizer.SpeakSsmlAsync(strSsml); if (speechSynthesisResult.Reason == ResultReason.SynthesizingAudioCompleted) { WaveFile waveFile = new WaveFile(speechSynthesisResult.AudioData); _processAndSave(waveFile, audioFilepaths); } else { new LogServices().AddTextToSpeechError(speechSynthesisResult, strText); } }
修改后,无论选择哪种音频输出格式,SpeakSsmlAsync都会直接返回语音数据,不会触发扬声器播放。
内容的提问来源于stack exchange,提问作者Richard
相关产品推荐
相关产品推荐

