You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PushStream流式传输Azure TTS时客户端响应大小为0的问题求助

问题描述

尝试用Fetch API和PassThrough将Azure TTS从服务器流式传输到客户端,预期分块接收音频流,但实际响应对象大小为0且无内容。调试后端日志显示已生成并发送流,但前端无论是直接读取Blob还是用ReadableStream,都无法获取有效内容。


后端语音生成函数

const generateSpeechFromText = async (text) => {
  const speechConfig = sdk.SpeechConfig.fromSubscription(
    process.env.SPEECH_KEY,
    process.env.SPEECH_REGION
  );
  speechConfig.speechSynthesisVoiceName = "en-US-JennyNeural";
  speechConfig.speechSynthesisOutputFormat =
    sdk.SpeechSynthesisOutputFormat.Audio16Khz32KBitRateMonoMp3;

  const synthesizer = new sdk.SpeechSynthesizer(speechConfig);

  return new Promise((resolve, reject) => {
    synthesizer.speakTextAsync(
      text,
      (result) => {
        if (result.reason === sdk.ResultReason.SynthesizingAudioCompleted) {
          const bufferStream = new PassThrough();
          bufferStream.end(Buffer.from(result.audioData));
          resolve(bufferStream);
        } else {
          console.error("Speech synthesis canceled: " + result.errorDetails);
          reject(new Error("Speech synthesis failed"));
        }
        synthesizer.close();
      },
      (error) => {
        console.error("Error in speech synthesis: " + error);
        synthesizer.close();
        reject(error);
      }
    );
  });
};

后端路由代码

app.get("/textToSpeech", async (request, reply) => {
  if (textWorks) {
    try {
      const stream = await generateSpeechFromText(
        textWorks
      );
      console.log("Stream created, sending to client: ", stream);
      reply.type("audio/mpeg").send(stream);
    } catch (err) {
      console.error(err);
      reply.status(500).send("Error in text-to-speech synthesis");
    }
  } else {
    reply.status(404).send("OpenAI response not found");
  }
});

前端请求代码

// Fetch TTS from Backend
export const fetchTTS = async (): Promise<Blob | null> => {
  try {
    const response = await fetch("http://localhost:3000/textToSpeech", {
      method: "GET",
    });
    // the response is size 0 and has no information
    if (!response.ok) {
      throw new Error(`HTTP error! status: ${response.status}`);
    }
    
    const body = response.body;
    console.log("Body", body);
    if (!body) {
      console.error("Response body is not a readable stream.");
      return null;
    }

    const reader = body.getReader();
    let chunks: Uint8Array[] = [];

    const read = async () => {
      const { done, value } = await reader.read();
      if (done) {
        return;
      }

      if (value) {
        chunks.push(value);
      }
      await read();
    };
    console.log("Chunks", chunks);

    await read();

    const audioBlob = new Blob(chunks, { type: "audio/mpeg" });

    // console.log("Audio Blob: ", audioBlob);
    // console.log("Audio Blob Size: ", audioBlob.size);

    return audioBlob.size > 0 ? audioBlob : null;
  } catch (error) {
    console.error("Error fetching text-to-speech audio:", error);
    return null;
  }
};

解决方案

1. 修复后端流式传输逻辑

当前后端是等音频完全合成后才一次性写入流,并非真正的分块传输。需监听Azure TTS的synthesizing事件,实时推送生成的音频块:

const generateSpeechFromText = async (text) => {
  const speechConfig = sdk.SpeechConfig.fromSubscription(
    process.env.SPEECH_KEY,
    process.env.SPEECH_REGION
  );
  speechConfig.speechSynthesisVoiceName = "en-US-JennyNeural";
  speechConfig.speechSynthesisOutputFormat =
    sdk.SpeechSynthesisOutputFormat.Audio16Khz32KBitRateMonoMp3;

  const synthesizer = new sdk.SpeechSynthesizer(speechConfig);
  const bufferStream = new PassThrough();

  return new Promise((resolve, reject) => {
    // 实时接收合成中的音频块
    synthesizer.synthesizing = (_, e) => {
      if (e.audioData) {
        bufferStream.write(Buffer.from(e.audioData));
      }
    };

    synthesizer.speakTextAsync(
      text,
      (result) => {
        if (result.reason === sdk.ResultReason.SynthesizingAudioCompleted) {
          bufferStream.end();
          resolve(bufferStream);
        } else {
          console.error("Speech synthesis canceled: " + result.errorDetails);
          bufferStream.destroy(new Error("Speech synthesis failed"));
          reject(new Error("Speech synthesis failed"));
        }
        synthesizer.close();
      },
      (error) => {
        console.error("Error in speech synthesis: " + error);
        bufferStream.destroy(error);
        synthesizer.close();
        reject(error);
      }
    );
  });
};

2. 明确启用分块传输头

在后端路由中显式设置分块传输头,确保客户端识别流式响应:

app.get("/textToSpeech", async (request, reply) => {
  if (textWorks) {
    try {
      const stream = await generateSpeechFromText(textWorks);
      console.log("Stream created, sending to client: ", stream);
      reply.header('Transfer-Encoding', 'chunked');
      reply.type("audio/mpeg").send(stream);
    } catch (err) {
      console.error(err);
      reply.status(500).send("Error in text-to-speech synthesis");
    }
  } else {
    reply.status(404).send("OpenAI response not found");
  }
});

3. 优化前端读取逻辑

调整日志时机,简化流读取循环,并支持实时处理音频块:

export const fetchTTS = async (): Promise<Blob | null> => {
  try {
    const response = await fetch("http://localhost:3000/textToSpeech", {
      method: "GET",
    });

    if (!response.ok) {
      throw new Error(`HTTP error! status: ${response.status}`);
    }
    
    const body = response.body;
    if (!body) {
      console.error("Response body is not a readable stream.");
      return null;
    }

    const reader = body.getReader();
    const chunks: Uint8Array[] = [];

    while (true) {
      const { done, value } = await reader.read();
      if (done) break;
      if (value) {
        chunks.push(value);
        // 可选:实时播放音频块
        // const blob = new Blob([value], { type: 'audio/mpeg' });
        // const url = URL.createObjectURL(blob);
        // audioElement.src = url;
      }
    }

    const audioBlob = new Blob(chunks, { type: "audio/mpeg" });
    return audioBlob.size > 0 ? audioBlob : null;
  } catch (error) {
    console.error("Error fetching text-to-speech audio:", error);
    return null;
  }
};

4. 排查跨域问题

若前后端域名不同,确保后端配置CORS允许前端请求:

// Fastify CORS配置示例
import cors from '@fastify/cors';
app.register(cors, {
  origin: "http://localhost:你的前端端口",
  methods: ["GET"]
});

内容的提问来源于stack exchange,提问作者Kaushik Verukonda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 14:24:51