Node.js结合Google Speech-to-Text转录音频结果为空求助
问题排查与解决方案
你的核心问题是Google Speech-to-Text返回的转录结果为空,结合代码来看,主要可能是音频格式不兼容、前端DOM元素未正确获取、文件读取逻辑缺陷这几个原因,以下是分步排查和修复方案:
1. 先解决前端显示问题(容易被忽略)
你的前端代码里未定义transcriptionDiv变量,导致即使后端返回了结果,前端也无法更新显示,会让你误以为转录结果为空。在<script>开头添加:
const transcriptionDiv = document.getElementById('transcription');
同时,把btnStop的事件监听改为只绑定一次,避免多次点击录制按钮后重复触发:
btnStop.addEventListener('click', async () => { // ...原stop逻辑 }, { once: true });
2. 核心问题:录制的音频格式不符合Google API要求
Google Speech-to-Text的LINEAR16编码要求音频是无压缩PCM格式的WAV文件,但浏览器默认的MediaRecorder输出的是WebM(Opus编码),哪怕你在Blob里指定type: 'audio/wav',也只是修改了MIME类型,实际编码还是不兼容,导致API无法解析音频内容。
方案一:前端直接录制符合要求的WAV音频
使用AudioContext捕获PCM原始数据,手动封装为WAV格式:
btnStart.addEventListener('click', async () => { const stream = await navigator.mediaDevices.getUserMedia({audio: true, video: false}); const audioContext = new AudioContext({ sampleRate: 44100 }); // 与API采样率保持一致 const mediaSource = audioContext.createMediaStreamSource(stream); const scriptProcessor = audioContext.createScriptProcessor(4096, 1, 1); // 单声道 let audioChunks = []; mediaSource.connect(scriptProcessor); scriptProcessor.connect(audioContext.destination); // 捕获PCM数据 scriptProcessor.onaudioprocess = (e) => { audioChunks.push(new Float32Array(e.inputBuffer.getChannelData(0))); }; btnStop.addEventListener('click', async () => { // 停止录制并清理资源 scriptProcessor.disconnect(); mediaSource.disconnect(); audioContext.close(); stream.getTracks().forEach(track => track.stop()); // 将Float32格式转换为16位PCM const pcmData = new Int16Array(audioChunks.reduce((acc, chunk) => acc + chunk.length, 0)); let offset = 0; for (const chunk of audioChunks) { for (let i = 0; i < chunk.length; i++) { const sample = Math.max(-1, Math.min(1, chunk[i])); pcmData[offset++] = sample < 0 ? sample * 0x8000 : sample * 0x7FFF; } } // 封装WAV文件头 const wavBlob = new Blob([ new Uint8Array([82,73,70,70]), // RIFF标识 new Uint32Array([44 + pcmData.length * 2]).buffer, // 文件总长度 new Uint8Array([87,65,86,69]), // WAVE标识 new Uint8Array([102,109,116,32]), // fmt子块 new Uint32Array([16]).buffer, // 子块长度 new Uint16Array([1]).buffer, // PCM编码 new Uint16Array([1]).buffer, // 单声道 new Uint32Array([44100]).buffer, // 采样率 new Uint32Array([44100 * 2]).buffer, // 字节率 new Uint16Array([2]).buffer, // 块对齐 new Uint16Array([16]).buffer, // 位深度 new Uint8Array([100,97,116,97]), // data子块 new Uint32Array([pcmData.length * 2]).buffer, // 数据长度 pcmData.buffer ], { type: 'audio/wav' }); // 上传到后端 const formData = new FormData(); formData.append('audio', wavBlob, 'recording.wav'); const response = await fetch('/transcribe', { method: 'POST', body: formData, }); const transcription = await response.text(); transcriptionDiv.innerHTML = `Transcription: ${transcription}`; }, { once: true }); });
方案二:后端转换音频格式(无需修改前端)
使用ffmpeg将上传的音频转换为LINEAR16格式:
- 安装依赖:
npm install fluent-ffmpeg
- 修改后端路由:
const ffmpeg = require('fluent-ffmpeg'); const path = require('path'); app.post('/transcribe', upload, async (req, res) => { try { const inputPath = req.file.path; const outputPath = path.join('uploads', 'converted.wav'); // 转换为LINEAR16编码的WAV await new Promise((resolve, reject) => { ffmpeg(inputPath) .output(outputPath) .audioCodec('pcm_s16le') // 对应LINEAR16 .audioChannels(1) .audioFrequency(44100) .on('end', resolve) .on('error', reject) .run(); }); const config = { encoding: 'LINEAR16', sampleRateHertz: 44100, languageCode: 'fr-FR', }; const audio = { content: fs.readFileSync(outputPath).toString('base64'), }; const request = { audio, config }; const [response] = await client.recognize(request); const transcription = response.results.map(result => result.alternatives[0].transcript).join('\n'); console.log('Transcription:', transcription); res.send(transcription); } catch (err) { console.error('Transcription error:', err); res.status(500).send('Transcription failed'); } });
3. 修复文件读取的潜在问题
当前multer配置固定文件名test-speech.wav,多次上传会覆盖旧文件,导致读取到不完整的音频。改为使用req.file.path直接读取上传的文件:
// 后端路由中替换audio.content的读取逻辑 const audio = { content: fs.readFileSync(req.file.path).toString('base64'), };
4. 验证Google API配置
确保:
- 已启用Google Cloud Speech-to-Text API
- 服务账号密钥配置正确(
client实例初始化无误) - 可以用一个已知有效的WAV文件测试API,确认API本身能返回结果
内容的提问来源于stack exchange,提问作者Alpha Col
相关产品推荐
相关产品推荐

