Azure Text-To-Speech SDK实现流式输入输出的技术问询
实现Azure TTS流式输入输出(边输入边播放)
当然可以用Azure Text-To-Speech SDK实现流式输入输出,实时生成并播放音频,无需等待全部文本处理完成。下面是具体的实现方案和代码示例:
核心思路
Azure TTS SDK支持流式音频输出和增量文本提交,结合这两个特性可以实现:
- 边输入文本(比如用户打字过程中逐步提交),边合成音频片段
- 实时播放收到的音频数据,延迟远低于全量生成后再播放的模式
- 无需将音频保存到本地文件,直接通过浏览器的
AudioContext实时解码播放
修改后的实现代码
// 初始化音频播放上下文与队列管理 let audioContext; let audioBufferQueue = []; let isPlaying = false; // 解码并播放单段音频缓冲区 async function playAudioBuffer(buffer) { if (!audioContext) { audioContext = new (window.AudioContext || window.webkitAudioContext)(); } try { const audioBuffer = await audioContext.decodeAudioData(buffer); audioBufferQueue.push(audioBuffer); if (!isPlaying) { processAudioQueue(); } } catch (err) { console.error("音频解码失败:", err); } } // 处理音频播放队列,确保片段连续播放 function processAudioQueue() { if (audioBufferQueue.length === 0) { isPlaying = false; return; } isPlaying = true; const buffer = audioBufferQueue.shift(); const source = audioContext.createBufferSource(); source.buffer = buffer; source.connect(audioContext.destination); source.onended = processAudioQueue; source.start(); } // 流式TTS封装类 class StreamingTTS { constructor(subscriptionKey, region) { // 初始化Speech配置 this.speechConfig = sdk.SpeechConfig.fromSubscription(subscriptionKey, region); // 选择低延迟的PCM格式,便于实时播放 this.speechConfig.speechSynthesisOutputFormat = sdk.SpeechSynthesisOutputFormat.Raw16Khz16BitMonoPcm; // 创建流式音频输出流,用于接收实时音频数据 this.audioOutputStream = sdk.PushAudioOutputStream.create(); this.audioConfig = sdk.AudioConfig.fromStreamOutput(this.audioOutputStream); // 初始化合成器 this.synthesizer = new sdk.SpeechSynthesizer(this.speechConfig, this.audioConfig); // 监听音频数据推送事件 this.audioOutputStream.on("data", (audioChunk) => { playAudioBuffer(audioChunk); }); // 监听合成完成/错误事件 this.synthesizer.synthesisCompleted = () => { this.audioOutputStream.close(); }; this.synthesizer.synthesisCanceled = (_, event) => { console.error("合成异常:", event.errorDetails); this.audioOutputStream.close(); }; } // 启动合成(提交初始文本) async start(initialText) { await this.synthesizer.startSpeakingTextAsync(initialText); } // 追加文本继续合成 async append(textChunk) { await this.synthesizer.sendTextAsync(textChunk); } // 停止合成并清理资源 async stop() { await this.synthesizer.stopSpeakingAsync(); this.synthesizer.close(); } } // 使用示例 // 替换为你的Azure订阅密钥和区域 const YOUR_KEY = "your-subscription-key"; const YOUR_REGION = "your-region"; const streamingTTS = new StreamingTTS(YOUR_KEY, YOUR_REGION); // 模拟用户逐步输入的场景 streamingTTS.start("你好,"); // 1秒后追加文本 setTimeout(() => { streamingTTS.append("这是Azure流式语音合成的演示。"); }, 1000); // 3秒后停止合成 setTimeout(() => { streamingTTS.stop(); }, 3000);
关键改进点说明
- 音频格式选择:放弃MP3改用
Raw16Khz16BitMonoPcm,避免压缩格式的解码延迟,更适合实时播放。 - 流式音频接收:通过
PushAudioOutputStream捕获SDK返回的每一段音频数据,无需保存到文件。 - 音频播放队列:维护播放队列确保音频片段连续播放,避免卡顿。
- 增量文本提交:通过
startSpeakingTextAsync启动合成,后续用sendTextAsync追加文本,实现边输入边合成。
注意事项
- 浏览器环境下需确保
AudioContext的兼容性(已做webkitAudioContext兼容处理)。 - 若在前端使用,需配置Azure TTS服务的CORS规则,允许你的域名访问。
- 合成结束后务必调用
stop()清理资源,避免内存泄漏。
内容的提问来源于stack exchange,提问作者JDLR
相关产品推荐
相关产品推荐

