Azure Text-to-speech返回字节数组到前端触发DOMException解码错误如何解决
问题根因
解码失败主要来自三个常见配置遗漏,不需要修改核心逻辑,补全对应配置即可解决:
- 后端未显式指定浏览器兼容的音频输出格式,且接口响应未正确设置Content-Type头
- 前端AudioContext的AudioBufferSourceNode未正确初始化,且decodeAudioData的异步调用未做播放触发处理
后端修改步骤
- 在SpeechConfig初始化后显式指定浏览器通用兼容的输出格式,优先选择
Riff16Khz16BitMonoPcm,该格式所有现代浏览器均原生支持:
public byte[] ConvertTextToSpeechStream(string text) { SpeechConfig speechConfiguration = SpeechConfig.FromEndpoint(new System.Uri("https://eastus.api.cognitive.microsoft.com/sts/v1.0/issuetoken"), "<subscription-key>"); // 新增这一行指定输出格式 speechConfiguration.SetSpeechSynthesisOutputFormat(SpeechSynthesisOutputFormat.Riff16Khz16BitMonoPcm); using (SpeechSynthesizer speechSynthesizer = new SpeechSynthesizer(speechConfiguration, null)) { SpeechSynthesisResult result = speechSynthesizer.SpeakTextAsync(text).Result; return result.AudioData; } }
- REST接口返回byte[]时,必须显式设置响应Content-Type为
audio/wav,避免框架默认返回非音频类型标识,.NET Core接口示例写法:
[HttpPost("tts")] public IActionResult GetSpeech([FromBody] TtsRequest request) { var audioBytes = ConvertTextToSpeechStream(request.Text); return File(audioBytes, "audio/wav"); }
前端修改步骤
原有JS代码中source变量未初始化,也缺少播放触发逻辑,调整后代码如下:
audioPlayByteStream: async function () { window.AudioContext = window.AudioContext || window.webkitAudioContext; const audioContext = new AudioContext(); // 每次播放需要新建独立的AudioBufferSourceNode,不可复用 const source = audioContext.createBufferSource(); try { const response = await fetch('my-endpoint-url', { method: "post", headers: { "Content-type": "application/json;charset=UTF-8", }, body: JSON.stringify({ Text: "This is my test text!" }), }); const buffer = await response.arrayBuffer(); const decodedData = await audioContext.decodeAudioData(buffer); source.buffer = decodedData; source.connect(audioContext.destination); // 触发音频播放 source.start(0); } catch (error) { console.log("请求或播放失败", error); } }
如果调整后仍有解码问题,可先将后端返回的byte[]保存为本地wav文件,用系统播放器测试播放,排除后端生成音频本身的异常。
内容的提问来源于stack exchange,提问作者Dejan Stamenov
相关产品推荐
相关产品推荐

