WebM Opus音频流式传输报错CHUNK_DEMUXER_ERROR_APPEND_FAILED求助
问题原因分析
- 数据类型错误:前端将WebM字节流(Uint8Array)转换为Uint16Array后追加到SourceBuffer,破坏了原始二进制结构,导致解析失败。WebM是基于字节的格式,必须直接传递Uint8Array。
- SourceBuffer重置不彻底:当前重置流程中,调用
sourceBuffer.remove(0, Infinity)后立即设置updating=false,但remove是异步操作,此时SourceBuffer可能仍处于更新状态,后续追加数据会触发冲突。 - 未处理剩余音频数据:收到END_AI_RESPONSE信号时,未将剩余的buffer数据追加到SourceBuffer,既丢失了最后一段音频,也可能导致SourceBuffer残留不完整数据。
解决方案
1. 修复数据类型传递错误
移除前端中new Uint16Array(buffer.slice(0, 512).buffer)的转换逻辑,直接将Uint8Array数据推入pendingAppendData。
2. 完善SourceBuffer异步重置流程
收到START_AI_RESPONSE信号时,等待remove操作完成后再重置状态,确保SourceBuffer完全清空后再接收新数据。
3. 处理剩余音频数据
在END_AI_RESPONSE时,将剩余buffer追加到SourceBuffer,保证音频完整性。
修改后的前端代码
///////////////////////////// Speech response ///////////////////////////// // Create a new MediaSource instance var mediaSource = new MediaSource(); // Create an audio element const audioElement = document.createElement('audio'); var updating = false; var pendingAppendData = []; var sourceBuffer; let buffer = new Uint8Array([]); async function websocketMessageHandler(event) { if (event.data.maxByteLength !== 1) { // Append new data to the buffer buffer = new Uint8Array([...buffer, ...new Uint8Array(event.data)]); if (buffer.length >= 512) { // 直接传递Uint8Array,不转换为Uint16Array pendingAppendData.push(buffer.slice(0, 512)); buffer = buffer.slice(512); appendData(); // Append the new data to the SourceBuffer if (audioElement.paused) audioElement.play(); } } else { const controlByte = new Uint8Array(event.data)[0]; if (controlByte === ControlMessagesEnum.END_AI_RESPONSE[0]) { // 追加剩余的buffer数据 if (buffer.length > 0) { pendingAppendData.push(buffer); buffer = new Uint8Array([]); appendData(); } } else if (controlByte === ControlMessagesEnum.START_AI_RESPONSE[0]) { if (audioElement.currentTime !== 0 && sourceBuffer) { updating = true; audioElement.pause(); audioElement.currentTime = 0; sourceBuffer.abort(); // 等待remove操作完成后再重置状态 sourceBuffer.remove(0, Infinity); sourceBuffer.addEventListener('updateend', () => { buffer = new Uint8Array([]); pendingAppendData = []; updating = false; }, { once: true }); } } } } // # add on message to ws. function addOnWebsocketMessageCallback(external_function){ // change websocket callback to include external function. async function on_websocket_message (event){ external_function(event); await websocketMessageHandler(event); } ws.onmessage = on_websocket_message; } // Set the MediaSource object as the source for the audio element audioElement.src = URL.createObjectURL(mediaSource); audioElement.onerror = (err)=>console.error(err.target.error); // Listen for the 'sourceopen' event to create and initialize the SourceBuffer mediaSource.addEventListener('sourceopen', function(){ console.log('source is open') if (!sourceBuffer) { sourceBuffer = mediaSource.addSourceBuffer('audio/webm; codecs="opus"'); sourceBuffer.mode = 'sequence'; this.onsourceclose = ()=>console.log('them don close me oooo'); } }); // Function to append data to the SourceBuffer function appendData(){ if (updating || pendingAppendData.length === 0) { return; } updating = true; const data = pendingAppendData.shift(); sourceBuffer.addEventListener('updateend', ()=>{ updating = false; appendData(); // Recursively call appendData to process the next chunk }, {once: true}); sourceBuffer.appendBuffer(data); }
后端注意事项
确保每次调用speech_synthesizer.start_speaking后,发送的第一个音频块包含完整的WebM EBML头部和轨道信息(Tracks元素)。部分TTS SDK在流式输出时,会在每次新合成开始时自动包含完整头部,若你的SDK没有此行为,需要手动确保第一个块包含完整的WebM容器结构,否则新的音频流无法被SourceBuffer正确解析。
内容的提问来源于stack exchange,提问作者Jonathan
相关产品推荐
相关产品推荐

