You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过WebSocket解析Azure TTS流式文本输入响应并生成可播放音频?

解决Azure语音服务流式音频块拼接问题

你遇到的问题大概率是没有正确分离响应块中的HTTP头部和音频二进制内容,或是没处理HTTP分块编码的结构,导致写入的文件混入非音频数据,播放器无法识别。以下是具体处理方案:

核心问题分析

Azure流式TTS的响应采用分块传输,每个Path:audio块的结构为:

HTTP响应头行(如X-RequestId、Content-Type等)
空行(\r\n\r\n)
二进制音频数据

如果直接把整个块内容写入文件,会把头部文本混入音频数据,导致文件损坏。另外HTTP分块编码的响应每个块开头还有十六进制长度标识,也需要跳过。

具体处理步骤

  1. 解析响应块,分离头部与音频数据
    逐块读取响应数据,找到空行(\r\n\r\n)的位置,分割出头部文本和二进制音频内容,只提取空行后的二进制数据,忽略头部;同时跳过Path:response的JSON块,仅处理Path:audio的块。

  2. 按顺序拼接音频数据
    将所有提取到的音频二进制数据按响应顺序依次追加到文件或音频缓冲区,确保顺序不颠倒。

  3. 保存为正确格式的文件
    由于Content-Type是audio/mpeg,保存时使用.mp3后缀,确保播放器能识别格式。

Node.js环境代码示例

const fs = require('fs');
const http = require('http');

// 创建写入流,保存为mp3文件
const audioFile = fs.createWriteStream('output.mp3');
const speechEndpoint = '你的语音服务端点';
const subscriptionKey = '你的订阅密钥';

const requestOptions = {
  hostname: new URL(speechEndpoint).hostname,
  path: '/api/texttospeech/v3.0/streaming',
  method: 'POST',
  headers: {
    'Content-Type': 'application/json',
    'Ocp-Apim-Subscription-Key': subscriptionKey,
    'Transfer-Encoding': 'chunked'
  }
};

const req = http.request(requestOptions, (res) => {
  let tempBuffer = Buffer.from([]);

  res.on('data', (chunk) => {
    tempBuffer = Buffer.concat([tempBuffer, chunk]);
    // 查找头部与音频的分隔符:空行
    const separatorPos = tempBuffer.indexOf('\r\n\r\n');
    
    if (separatorPos !== -1) {
      const headerStr = tempBuffer.slice(0, separatorPos).toString('utf8');
      // 仅处理Path为audio的块
      if (headerStr.includes('Path:audio')) {
        const audioChunk = tempBuffer.slice(separatorPos + 4);
        audioFile.write(audioChunk);
      }
      // 重置临时缓冲区,处理剩余数据
      tempBuffer = Buffer.from([]);
    }
  });

  res.on('end', () => {
    audioFile.end();
    console.log('音频文件生成完成');
  });
});

// 发送TTS请求体(实际可扩展为流式文本输入)
req.write(JSON.stringify({
  voice: { name: 'zh-CN-XiaoxiaoNeural' },
  input: { text: '这是测试的流式语音合成内容' }
}));
req.end();

浏览器环境处理示例

async function streamTTS() {
  const response = await fetch('你的语音服务端点/api/texttospeech/v3.0/streaming', {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      'Ocp-Apim-Subscription-Key': '你的订阅密钥'
    },
    body: JSON.stringify({
      voice: { name: 'zh-CN-XiaoxiaoNeural' },
      input: { text: '浏览器流式测试' }
    })
  });

  const reader = response.body.getReader();
  const audioChunks = [];
  let tempBuffer = new Uint8Array();

  while (true) {
    const { done, value } = await reader.read();
    if (done) break;

    // 合并临时缓冲区与当前块
    const combined = new Uint8Array(tempBuffer.length + value.length);
    combined.set(tempBuffer);
    combined.set(value, tempBuffer.length);
    tempBuffer = combined;

    // 查找空行分隔符
    const separator = new TextEncoder().encode('\r\n\r\n');
    let separatorPos = -1;
    for (let i = 0; i <= tempBuffer.length - separator.length; i++) {
      if (tempBuffer.slice(i, i + separator.length).every((val, idx) => val === separator[idx])) {
        separatorPos = i;
        break;
      }
    }

    if (separatorPos !== -1) {
      const headerStr = new TextDecoder().decode(tempBuffer.slice(0, separatorPos));
      if (headerStr.includes('Path:audio')) {
        const audioData = tempBuffer.slice(separatorPos + separator.length);
        audioChunks.push(audioData);
      }
      tempBuffer = tempBuffer.slice(separatorPos + separator.length);
    }
  }

  // 生成下载链接
  const audioBlob = new Blob(audioChunks, { type: 'audio/mpeg' });
  const url = URL.createObjectURL(audioBlob);
  const a = document.createElement('a');
  a.href = url;
  a.download = 'output.mp3';
  a.click();
  URL.revokeObjectURL(url);
}

内容的提问来源于stack exchange,提问作者xiaoyu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 12:44:58