使用PushStream流式传输Azure TTS时客户端响应大小为0的问题求助
问题描述
尝试用Fetch API和PassThrough将Azure TTS从服务器流式传输到客户端,预期分块接收音频流,但实际响应对象大小为0且无内容。调试后端日志显示已生成并发送流,但前端无论是直接读取Blob还是用ReadableStream,都无法获取有效内容。
后端语音生成函数
const generateSpeechFromText = async (text) => { const speechConfig = sdk.SpeechConfig.fromSubscription( process.env.SPEECH_KEY, process.env.SPEECH_REGION ); speechConfig.speechSynthesisVoiceName = "en-US-JennyNeural"; speechConfig.speechSynthesisOutputFormat = sdk.SpeechSynthesisOutputFormat.Audio16Khz32KBitRateMonoMp3; const synthesizer = new sdk.SpeechSynthesizer(speechConfig); return new Promise((resolve, reject) => { synthesizer.speakTextAsync( text, (result) => { if (result.reason === sdk.ResultReason.SynthesizingAudioCompleted) { const bufferStream = new PassThrough(); bufferStream.end(Buffer.from(result.audioData)); resolve(bufferStream); } else { console.error("Speech synthesis canceled: " + result.errorDetails); reject(new Error("Speech synthesis failed")); } synthesizer.close(); }, (error) => { console.error("Error in speech synthesis: " + error); synthesizer.close(); reject(error); } ); }); };
后端路由代码
app.get("/textToSpeech", async (request, reply) => { if (textWorks) { try { const stream = await generateSpeechFromText( textWorks ); console.log("Stream created, sending to client: ", stream); reply.type("audio/mpeg").send(stream); } catch (err) { console.error(err); reply.status(500).send("Error in text-to-speech synthesis"); } } else { reply.status(404).send("OpenAI response not found"); } });
前端请求代码
// Fetch TTS from Backend export const fetchTTS = async (): Promise<Blob | null> => { try { const response = await fetch("http://localhost:3000/textToSpeech", { method: "GET", }); // the response is size 0 and has no information if (!response.ok) { throw new Error(`HTTP error! status: ${response.status}`); } const body = response.body; console.log("Body", body); if (!body) { console.error("Response body is not a readable stream."); return null; } const reader = body.getReader(); let chunks: Uint8Array[] = []; const read = async () => { const { done, value } = await reader.read(); if (done) { return; } if (value) { chunks.push(value); } await read(); }; console.log("Chunks", chunks); await read(); const audioBlob = new Blob(chunks, { type: "audio/mpeg" }); // console.log("Audio Blob: ", audioBlob); // console.log("Audio Blob Size: ", audioBlob.size); return audioBlob.size > 0 ? audioBlob : null; } catch (error) { console.error("Error fetching text-to-speech audio:", error); return null; } };
解决方案
1. 修复后端流式传输逻辑
当前后端是等音频完全合成后才一次性写入流,并非真正的分块传输。需监听Azure TTS的synthesizing事件,实时推送生成的音频块:
const generateSpeechFromText = async (text) => { const speechConfig = sdk.SpeechConfig.fromSubscription( process.env.SPEECH_KEY, process.env.SPEECH_REGION ); speechConfig.speechSynthesisVoiceName = "en-US-JennyNeural"; speechConfig.speechSynthesisOutputFormat = sdk.SpeechSynthesisOutputFormat.Audio16Khz32KBitRateMonoMp3; const synthesizer = new sdk.SpeechSynthesizer(speechConfig); const bufferStream = new PassThrough(); return new Promise((resolve, reject) => { // 实时接收合成中的音频块 synthesizer.synthesizing = (_, e) => { if (e.audioData) { bufferStream.write(Buffer.from(e.audioData)); } }; synthesizer.speakTextAsync( text, (result) => { if (result.reason === sdk.ResultReason.SynthesizingAudioCompleted) { bufferStream.end(); resolve(bufferStream); } else { console.error("Speech synthesis canceled: " + result.errorDetails); bufferStream.destroy(new Error("Speech synthesis failed")); reject(new Error("Speech synthesis failed")); } synthesizer.close(); }, (error) => { console.error("Error in speech synthesis: " + error); bufferStream.destroy(error); synthesizer.close(); reject(error); } ); }); };
2. 明确启用分块传输头
在后端路由中显式设置分块传输头,确保客户端识别流式响应:
app.get("/textToSpeech", async (request, reply) => { if (textWorks) { try { const stream = await generateSpeechFromText(textWorks); console.log("Stream created, sending to client: ", stream); reply.header('Transfer-Encoding', 'chunked'); reply.type("audio/mpeg").send(stream); } catch (err) { console.error(err); reply.status(500).send("Error in text-to-speech synthesis"); } } else { reply.status(404).send("OpenAI response not found"); } });
3. 优化前端读取逻辑
调整日志时机,简化流读取循环,并支持实时处理音频块:
export const fetchTTS = async (): Promise<Blob | null> => { try { const response = await fetch("http://localhost:3000/textToSpeech", { method: "GET", }); if (!response.ok) { throw new Error(`HTTP error! status: ${response.status}`); } const body = response.body; if (!body) { console.error("Response body is not a readable stream."); return null; } const reader = body.getReader(); const chunks: Uint8Array[] = []; while (true) { const { done, value } = await reader.read(); if (done) break; if (value) { chunks.push(value); // 可选:实时播放音频块 // const blob = new Blob([value], { type: 'audio/mpeg' }); // const url = URL.createObjectURL(blob); // audioElement.src = url; } } const audioBlob = new Blob(chunks, { type: "audio/mpeg" }); return audioBlob.size > 0 ? audioBlob : null; } catch (error) { console.error("Error fetching text-to-speech audio:", error); return null; } };
4. 排查跨域问题
若前后端域名不同,确保后端配置CORS允许前端请求:
// Fastify CORS配置示例 import cors from '@fastify/cors'; app.register(cors, { origin: "http://localhost:你的前端端口", methods: ["GET"] });
内容的提问来源于stack exchange,提问作者Kaushik Verukonda
相关产品推荐
相关产品推荐

