如何将流式ChatGPT API(createChatCompletion)响应转换为语音
把ChatGPT流式响应实时转语音的实现方法
核心逻辑
你已经能流式获取ChatGPT的文本响应,要实现实时语音播报,核心就是每收到一段文本片段,立刻传给语音合成工具,不用等全部内容加载完成。下面是基于浏览器原生SpeechSynthesis API的实现方案,无需额外依赖。
修改后的完整代码
// 初始化语音合成实例(放在全局或函数外,避免重复创建) const synth = window.speechSynthesis; let currentUtterances = []; // 存储当前播报实例,方便中途终止 async function gpt(input) { // 清空之前的语音播报队列,避免新旧语音重叠 synth.cancel(); currentUtterances = []; const prompt = `The user has asked the following question: ${input} .Please respond to the user in two to three sentences in very simple way.`; const response = await fetch("/api/chat", { method: "POST", headers: { "Content-Type": "application/json", }, body: JSON.stringify({ prompt, }), }); if (!response.ok) { throw new Error(response.statusText); } const data = response.body; setIsListening(false); setGptloading(false); if (!data) { return; } const reader = data.getReader(); const decoder = new TextDecoder("utf-8"); let done = false; while (!done) { const { value, done: doneReading } = await reader.read(); done = doneReading; const chunkValue = decoder.decode(value); // 更新页面展示文本 setText((prev) => prev + chunkValue); // 实时合成当前文本片段的语音 if (chunkValue.trim()) { // 过滤空白片段,避免无效播报 const utterance = new SpeechSynthesisUtterance(chunkValue); // 配置语音参数:中文播报、语速、音量等 utterance.lang = "zh-CN"; utterance.rate = 1; // 语速范围0.1-10,默认1 utterance.volume = 1; // 音量范围0-1,默认1 currentUtterances.push(utterance); synth.speak(utterance); } } } // 可选:中途终止语音播报的函数 function stopSpeech() { synth.cancel(); currentUtterances = []; }
关键细节说明
- 语音队列管理:每次调用
gpt前执行synth.cancel(),清空旧的播报队列,防止语音重叠。 - 无效片段过滤:通过
chunkValue.trim()判断,避免空文本或空白字符触发不必要的播报。 - 语音参数自定义:可根据需求调整
utterance.lang(比如改成en-US实现英文播报)、语速、音量等参数。 - 中途终止功能:提供
stopSpeech函数,方便用户随时停止当前语音播报。
进阶方案(第三方TTS)
如果需要更自然的语音效果,可以用国内云厂商的流式TTS服务(如百度智能云、阿里云语音合成),核心逻辑类似:
- 每拿到一段文本片段,立刻调用TTS的流式接口
- 将返回的音频片段通过
AudioContext实时播放
内容的提问来源于stack exchange,提问作者RAHUL KUMAR
相关产品推荐
相关产品推荐

