You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将流式ChatGPT API(createChatCompletion)响应转换为语音

把ChatGPT流式响应实时转语音的实现方法

核心逻辑

你已经能流式获取ChatGPT的文本响应,要实现实时语音播报,核心就是每收到一段文本片段,立刻传给语音合成工具,不用等全部内容加载完成。下面是基于浏览器原生SpeechSynthesis API的实现方案,无需额外依赖。

修改后的完整代码

// 初始化语音合成实例(放在全局或函数外,避免重复创建)
const synth = window.speechSynthesis;
let currentUtterances = []; // 存储当前播报实例,方便中途终止

async function gpt(input) {
  // 清空之前的语音播报队列,避免新旧语音重叠
  synth.cancel();
  currentUtterances = [];

  const prompt = `The user has asked the following question: ${input} .Please respond to the user  in two to three sentences in very simple way.`;
  const response = await fetch("/api/chat", {
    method: "POST",
    headers: {
      "Content-Type": "application/json",
    },
    body: JSON.stringify({
      prompt,
    }),
  });

  if (!response.ok) {
    throw new Error(response.statusText);
  }

  const data = response.body;
  setIsListening(false);
  setGptloading(false);

  if (!data) {
    return;
  }

  const reader = data.getReader();
  const decoder = new TextDecoder("utf-8");
  let done = false;

  while (!done) {
    const { value, done: doneReading } = await reader.read();
    done = doneReading;
    const chunkValue = decoder.decode(value);
    
    // 更新页面展示文本
    setText((prev) => prev + chunkValue);

    // 实时合成当前文本片段的语音
    if (chunkValue.trim()) { // 过滤空白片段,避免无效播报
      const utterance = new SpeechSynthesisUtterance(chunkValue);
      // 配置语音参数:中文播报、语速、音量等
      utterance.lang = "zh-CN";
      utterance.rate = 1; // 语速范围0.1-10,默认1
      utterance.volume = 1; // 音量范围0-1,默认1
      
      currentUtterances.push(utterance);
      synth.speak(utterance);
    }
  }
}

// 可选:中途终止语音播报的函数
function stopSpeech() {
  synth.cancel();
  currentUtterances = [];
}

关键细节说明

  • 语音队列管理:每次调用gpt前执行synth.cancel(),清空旧的播报队列,防止语音重叠。
  • 无效片段过滤:通过chunkValue.trim()判断,避免空文本或空白字符触发不必要的播报。
  • 语音参数自定义:可根据需求调整utterance.lang(比如改成en-US实现英文播报)、语速、音量等参数。
  • 中途终止功能:提供stopSpeech函数,方便用户随时停止当前语音播报。

进阶方案(第三方TTS)

如果需要更自然的语音效果,可以用国内云厂商的流式TTS服务(如百度智能云、阿里云语音合成),核心逻辑类似:

  • 每拿到一段文本片段,立刻调用TTS的流式接口
  • 将返回的音频片段通过AudioContext实时播放

内容的提问来源于stack exchange,提问作者RAHUL KUMAR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 18:44:26