You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Node.js结合Google Speech-to-Text转录音频结果为空求助

问题排查与解决方案

你的核心问题是Google Speech-to-Text返回的转录结果为空,结合代码来看,主要可能是音频格式不兼容、前端DOM元素未正确获取、文件读取逻辑缺陷这几个原因,以下是分步排查和修复方案:


1. 先解决前端显示问题(容易被忽略)

你的前端代码里未定义transcriptionDiv变量,导致即使后端返回了结果,前端也无法更新显示,会让你误以为转录结果为空。在<script>开头添加:

const transcriptionDiv = document.getElementById('transcription');

同时,把btnStop的事件监听改为只绑定一次,避免多次点击录制按钮后重复触发:

btnStop.addEventListener('click', async () => {
    // ...原stop逻辑
}, { once: true });

2. 核心问题:录制的音频格式不符合Google API要求

Google Speech-to-Text的LINEAR16编码要求音频是无压缩PCM格式的WAV文件,但浏览器默认的MediaRecorder输出的是WebM(Opus编码),哪怕你在Blob里指定type: 'audio/wav',也只是修改了MIME类型,实际编码还是不兼容,导致API无法解析音频内容。

方案一:前端直接录制符合要求的WAV音频

使用AudioContext捕获PCM原始数据,手动封装为WAV格式:

btnStart.addEventListener('click', async () => {
    const stream = await navigator.mediaDevices.getUserMedia({audio: true, video: false});
    const audioContext = new AudioContext({ sampleRate: 44100 }); // 与API采样率保持一致
    const mediaSource = audioContext.createMediaStreamSource(stream);
    const scriptProcessor = audioContext.createScriptProcessor(4096, 1, 1); // 单声道

    let audioChunks = [];

    mediaSource.connect(scriptProcessor);
    scriptProcessor.connect(audioContext.destination);

    // 捕获PCM数据
    scriptProcessor.onaudioprocess = (e) => {
        audioChunks.push(new Float32Array(e.inputBuffer.getChannelData(0)));
    };

    btnStop.addEventListener('click', async () => {
        // 停止录制并清理资源
        scriptProcessor.disconnect();
        mediaSource.disconnect();
        audioContext.close();
        stream.getTracks().forEach(track => track.stop());

        // 将Float32格式转换为16位PCM
        const pcmData = new Int16Array(audioChunks.reduce((acc, chunk) => acc + chunk.length, 0));
        let offset = 0;
        for (const chunk of audioChunks) {
            for (let i = 0; i < chunk.length; i++) {
                const sample = Math.max(-1, Math.min(1, chunk[i]));
                pcmData[offset++] = sample < 0 ? sample * 0x8000 : sample * 0x7FFF;
            }
        }

        // 封装WAV文件头
        const wavBlob = new Blob([
            new Uint8Array([82,73,70,70]), // RIFF标识
            new Uint32Array([44 + pcmData.length * 2]).buffer, // 文件总长度
            new Uint8Array([87,65,86,69]), // WAVE标识
            new Uint8Array([102,109,116,32]), // fmt子块
            new Uint32Array([16]).buffer, // 子块长度
            new Uint16Array([1]).buffer, // PCM编码
            new Uint16Array([1]).buffer, // 单声道
            new Uint32Array([44100]).buffer, // 采样率
            new Uint32Array([44100 * 2]).buffer, // 字节率
            new Uint16Array([2]).buffer, // 块对齐
            new Uint16Array([16]).buffer, // 位深度
            new Uint8Array([100,97,116,97]), // data子块
            new Uint32Array([pcmData.length * 2]).buffer, // 数据长度
            pcmData.buffer
        ], { type: 'audio/wav' });

        // 上传到后端
        const formData = new FormData();
        formData.append('audio', wavBlob, 'recording.wav');

        const response = await fetch('/transcribe', {
            method: 'POST',
            body: formData,
        });

        const transcription = await response.text();
        transcriptionDiv.innerHTML = `Transcription: ${transcription}`;
    }, { once: true });
});

方案二:后端转换音频格式(无需修改前端)

使用ffmpeg将上传的音频转换为LINEAR16格式:

  1. 安装依赖:
npm install fluent-ffmpeg
  1. 修改后端路由:
const ffmpeg = require('fluent-ffmpeg');
const path = require('path');

app.post('/transcribe', upload, async (req, res) => {
    try {
        const inputPath = req.file.path;
        const outputPath = path.join('uploads', 'converted.wav');

        // 转换为LINEAR16编码的WAV
        await new Promise((resolve, reject) => {
            ffmpeg(inputPath)
                .output(outputPath)
                .audioCodec('pcm_s16le') // 对应LINEAR16
                .audioChannels(1)
                .audioFrequency(44100)
                .on('end', resolve)
                .on('error', reject)
                .run();
        });

        const config = {
            encoding: 'LINEAR16',
            sampleRateHertz: 44100,
            languageCode: 'fr-FR', 
        };

        const audio = {
            content: fs.readFileSync(outputPath).toString('base64'),
        };

        const request = { audio, config };
        const [response] = await client.recognize(request);
        const transcription = response.results.map(result => result.alternatives[0].transcript).join('\n');

        console.log('Transcription:', transcription);
        res.send(transcription);

    } catch (err) {
        console.error('Transcription error:', err);
        res.status(500).send('Transcription failed');
    }
});

3. 修复文件读取的潜在问题

当前multer配置固定文件名test-speech.wav,多次上传会覆盖旧文件,导致读取到不完整的音频。改为使用req.file.path直接读取上传的文件:

// 后端路由中替换audio.content的读取逻辑
const audio = {
    content: fs.readFileSync(req.file.path).toString('base64'),
};

4. 验证Google API配置

确保:

  • 已启用Google Cloud Speech-to-Text API
  • 服务账号密钥配置正确(client实例初始化无误)
  • 可以用一个已知有效的WAV文件测试API,确认API本身能返回结果

内容的提问来源于stack exchange,提问作者Alpha Col

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 19:59:55