You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Speech-to-Text API的longrunningrecognize端点返回空结果求助

问题:Google Speech-to-Text API音频转写返回空结果

我在Cloud Function中调用Google Speech-to-Text API处理一段4分钟、约40MB的.wav格式音频,始终得到空结果,response.results返回空值。怀疑是配置问题,但查遍文档没找到解决线索。

我的代码:

exports.transcribeAudio = functions
  .runWith({ timeoutSeconds: 540 })
  .firestore.document("users/{userId}/transcriptions/{transcriptionId}")
  .onCreate(async (snap, context) => {
    const file = snap.data();
    const userId = context.params.userId;
    const userRef = admin.firestore().collection("users").doc(userId);
    const userDoc = await userRef.get();
    const credits = userDoc.data().credits;
    try {
      if (credits <= 0) {
        throw "Not enough credits";
      }
      const downloadUrl = new URL(file.url);
      const request = {
        config: {
          encoding: "LINEAR16",
          sampleRateHertz: 16000,
          languageCode: "en-US",
        },
        audio: {
          uri: downloadUrl,
        },
      };

      const [operation] = await speechClient.longRunningRecognize(request);
      const [response] = await operation.promise();

      const transcription = response.results
        .map((result) => result.alternatives[0].transcript)
        .join("\n");

      const transcriptionsRef = userRef
        .collection("transcriptions")
        .doc(context.params.transcriptionId);
      await transcriptionsRef.update({ text: transcription });
      const newCredits = credits - 1;
      userRef.update({ credits: newCredits });
    } catch (error) {
      console.log(error);
    }
  });

可能的解决方向:

  • 检查音频URI的访问权限:Speech-to-Text API必须能读取你提供的音频文件。如果是Cloud Storage的文件,直接用gs://bucket-name/file-path格式的URI最可靠,避免HTTP URL的权限问题。如果用HTTP URL,必须确保该链接是公开可访问的,或者API的服务账号有权限访问这个资源。
  • 核对音频参数与配置一致性:代码里配置的encoding: "LINEAR16"和sampleRateHertz: 16000必须和实际音频文件的参数完全匹配。可以用ffmpeg -i your-audio.wav命令查看音频的真实编码、采样率、声道数,如果是立体声,需要在config里添加audioChannelCount: 2。
  • 验证音频内容有效性:如果音频本身无声、噪音过大,或者实际语言和配置的languageCode: "en-US"不匹配,API会返回空结果。建议截取一段清晰的音频片段测试,确认内容可被识别。
  • 完善错误日志排查:当前catch块仅打印错误,建议补充打印完整的response对象,看是否有隐藏的错误提示或状态信息。同时可以把错误信息写入Firestore,方便后续排查。
  • 调整长音频处理配置:虽然用了longRunningRecognize(适合长音频),但可以尝试添加enableWordTimeOffsets: true等可选参数,或者使用model: "video"(更适合长音频转写)。

代码修改示例:

exports.transcribeAudio = functions
  .runWith({ timeoutSeconds: 540 })
  .firestore.document("users/{userId}/transcriptions/{transcriptionId}")
  .onCreate(async (snap, context) => {
    const file = snap.data();
    const userId = context.params.userId;
    const userRef = admin.firestore().collection("users").doc(userId);
    const userDoc = await userRef.get();
    const credits = userDoc.data().credits;
    const transcriptionsRef = userRef
      .collection("transcriptions")
      .doc(context.params.transcriptionId);
    try {
      if (credits <= 0) {
        throw "Not enough credits";
      }
      // 优先使用Cloud Storage的gs路径,替换成你的桶名和文件路径逻辑
      const audioUri = file.url.startsWith('gs://') 
        ? file.url 
        : `gs://your-bucket-name/${file.filePath}`;
      
      const request = {
        config: {
          encoding: "LINEAR16",
          sampleRateHertz: 16000,
          languageCode: "en-US",
          audioChannelCount: 2, // 如果是立体声请开启
          model: "video", // 适合长音频转写
          enableAutomaticPunctuation: true
        },
        audio: {
          uri: audioUri,
        },
      };

      const [operation] = await speechClient.longRunningRecognize(request);
      const [response] = await operation.promise();

      // 打印完整response排查问题
      console.log("API响应:", response);

      if (!response.results || response.results.length === 0) {
        throw new Error("未识别到音频内容");
      }

      const transcription = response.results
        .map((result) => result.alternatives[0].transcript)
        .join("\n");

      await transcriptionsRef.update({ text: transcription });
      await userRef.update({ credits: credits - 1 });
    } catch (error) {
      console.error("转写失败:", error);
      // 将错误写入Firestore
      await transcriptionsRef.update({ 
        error: error.message || JSON.stringify(error),
        text: ""
      });
    }
  });

内容的提问来源于stack exchange,提问作者Harsh Vardhan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 20:34:54