请求Opus编码时Azure TTS生成乱码音频问题排查
问题分析与修复
你的代码核心问题是误解了Azure Speech SDK返回的Opus格式:Audio24Khz16Bit48KbpsMonoOpus是Ogg容器封装的Opus音频,而非裸Opus数据包。直接将整个回调chunk当作裸Opus包解码,会导致解析错误,出现音频乱码、缺失,同时因包含Ogg容器头数据,压缩比看起来远低于预期。
具体错误点
- 格式判断错误:Azure返回的音频数据是Ogg封装格式,包含容器头和页结构,不是可直接解码的裸Opus包,
opus_packet_get_nb_frames的断言能通过只是巧合,实际输入数据不符合Opus包结构。 - 解码逻辑错误:直接对Ogg页数据调用
opus_decode,解码器无法识别容器结构,解码出的PCM数据完全错误。
修复方案
方案一:直接获取PCM格式(最简单)
如果不需要处理Opus格式,直接让Azure返回原始PCM,跳过解码步骤:
修改输出格式为:
azure_speech_config->SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat::Audio24Khz16BitMonoPcm);
回调中无需解码,直接写入文件:
azure_synth->Synthesizing += [&in_bytes, &decoded_bytes, fp](const SpeechSynthesisEventArgs& e) { printf("Synthesizing event received with audio chunk of %zu bytes\n", e.Result->GetAudioData()->size()); auto audio_data = e.Result->GetAudioData(); in_bytes += audio_data->size(); fwrite(audio_data->data(), 1, audio_data->size(), fp); decoded_bytes += audio_data->size(); };
生成的out.raw可直接用aplay out.raw -f S16_LE -r 24000 -c 1播放,无任何问题。
方案二:解析Ogg容器后解码Opus(保留Opus处理流程)
如果必须处理Opus格式,需要引入libogg库解析Ogg容器,提取裸Opus包后再解码。以下是修改后的完整代码:
#include <stdio.h> #include <string> #include <assert.h> #include <vector> #include <speechapi_cxx.h> #include <opus.h> #include <ogg/ogg.h> using namespace Microsoft::CognitiveServices::Speech; static const std::string subscription_key = "abcd1234"; // 替换为有效密钥 static const std::string service_region = "westus"; static const std::string text = "Hi, this is Azure"; static const int sample_rate = 24000; #define MAX_FRAME_SIZE 6*960 // 来自Opus示例 int main(int argc, char **argv) { // 初始化Opus解码器 int err; OpusDecoder* opus_decoder = opus_decoder_create(sample_rate, 1, &err); assert(err == OPUS_OK); // 初始化Ogg同步和流状态 ogg_sync_state ogg_sync; ogg_stream_state ogg_stream; ogg_page ogg_page; ogg_packet ogg_packet; int ogg_stream_initialized = 0; ogg_sync_init(&ogg_sync); // 创建Azure客户端 auto azure_speech_config = SpeechConfig::FromSubscription(subscription_key, service_region); azure_speech_config->SetSpeechSynthesisVoiceName("en-US-JennyNeural"); azure_speech_config->SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat::Audio24Khz16Bit48KbpsMonoOpus); auto azure_synth = SpeechSynthesizer::FromConfig(azure_speech_config, NULL); FILE* fp = fopen("out.raw", "w"); int in_bytes=0, decoded_bytes=0; // 回调处理音频数据 azure_synth->Synthesizing += [&](const SpeechSynthesisEventArgs& e) { auto audio_data = e.Result->GetAudioData(); size_t data_size = audio_data->size(); printf("Synthesizing event received with audio chunk of %zu bytes\n", data_size); in_bytes += data_size; // 将数据写入Ogg同步缓冲区 char* ogg_buf = ogg_sync_buffer(&ogg_sync, data_size); memcpy(ogg_buf, audio_data->data(), data_size); ogg_sync_wrote(&ogg_sync, data_size); // 提取Ogg页 while (ogg_sync_pageout(&ogg_sync, &ogg_page) == 1) { // 初始化流状态(仅第一次遇到头页时) if (!ogg_stream_initialized) { if (ogg_stream_init(&ogg_stream, ogg_page.header->serialno) != 0) { fprintf(stderr, "Failed to initialize Ogg stream\n"); return; } ogg_stream_initialized = 1; } // 将页放入流 if (ogg_stream_pagein(&ogg_stream, &ogg_page) < 0) { fprintf(stderr, "Corrupted Ogg page\n"); continue; } // 提取Opus数据包 int packet_ret; while ((packet_ret = ogg_stream_packetout(&ogg_stream, &ogg_packet)) != 0) { if (packet_ret < 0) { fprintf(stderr, "Corrupted Ogg packet\n"); continue; } // 跳过Opus头包(ID头和评论头) if (ogg_packet.b_o_s) { // 验证是否为Opus头 if (memcmp(ogg_packet.packet, "OpusHead", 8) != 0) { fprintf(stderr, "Not an Opus Ogg stream\n"); return; } continue; } if (memcmp(ogg_packet.packet, "OpusTags", 8) == 0) { continue; } // 解码裸Opus包 std::vector<uint8_t> decoded_data(MAX_FRAME_SIZE); int decoded_samples = opus_decode(opus_decoder, ogg_packet.packet, ogg_packet.bytes, (opus_int16*)decoded_data.data(), MAX_FRAME_SIZE/sizeof(opus_int16), 0); assert(decoded_samples > 0); int decoded_byte_size = decoded_samples * sizeof(opus_int16); printf("Decoded to %d bytes\n", decoded_byte_size); fwrite(decoded_data.data(), 1, decoded_byte_size, fp); decoded_bytes += decoded_byte_size; } } }; // 执行TTS auto result = azure_synth->SpeakText(text); printf("Done, got %d bytes, decoded to %d bytes\n", in_bytes, decoded_bytes); // 清理资源 fclose(fp); opus_decoder_destroy(opus_decoder); if (ogg_stream_initialized) { ogg_stream_clear(&ogg_stream); } ogg_sync_clear(&ogg_sync); }
编译与运行
编译时需要链接libspeechapi、libopus和libogg:
g++ -o tts_opus_decode tts_opus_decode.cpp -lspeechapi -lopus -logg
运行后生成的out.raw可通过原命令播放,音频乱码和缺失问题会解决,压缩比也会恢复正常。
内容的提问来源于stack exchange,提问作者itaych
相关产品推荐
相关产品推荐

