You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Cloud Speech API桌面可用但iPhone Chrome移动端报错求助

Google Cloud Speech API在iPhone Chrome中返回400错误的解决方案

问题背景

Web服务接收用户音频并通过Google Cloud Speech API转写文本,桌面端(本地/服务器部署)运行正常,但在iPhone的Chrome浏览器中调用时返回:
google.api_core.exceptions.InvalidArgument: 400 RecognitionAudio not set.
移动端请求日志额外显示:
details = "Invalid recognition 'config': bad encoding.."
已确认服务端收到的音频数据非空,但移动端录制的音频存在编码适配问题。

前端录制代码(JavaScript)

// Store the audio data in chunks as it is recorded
mediaRecorder.addEventListener("dataavailable", (event) => {
  chunks.push(event.data);
});

// When recording stops, create a blob from the chunks and send it to the server
mediaRecorder.addEventListener("stop", () => {
  const blob = new Blob(chunks, { type: "audio/wav" });
  const url = URL.createObjectURL(blob);
  audio.src = url;
  uploadAudio(blob); // Send the audio data to the server
  chunks = []; // Clear the chunks array
});

服务端处理代码(Python)

audio_bytes = request.get_data()

# Instantiate a client
client = speech.SpeechClient()

# Settings for Speech v1:
# Initialize request arguments
config = speech.RecognitionConfig()
config.language_code = LANGUAGE
config.model = "latest_long"
config.enable_automatic_punctuation = True

audio = speech.RecognitionAudio()
audio.content = audio_bytes

speech_request = speech.RecognizeRequest(
    config=config,
    audio=audio,
)

# Make the request
response = client.recognize(request=speech_request)

解决方案

1. 强制指定Speech API的音频编码参数

尽管文档说明WAV无需设置编码,但移动端浏览器生成的音频格式可能和桌面端存在差异(比如iPhone Chrome默认用AAC编码而非PCM)。在服务端的RecognitionConfig中明确添加编码配置:

config.audio_channel_count = 1  # 适配移动端常见的单声道录制
config.sample_rate_hertz = 16000  # Speech API推荐的16kHz采样率
config.encoding = speech.RecognitionConfig.AudioEncoding.LINEAR16  # WAV对应的PCM编码

2. 修正前端MediaRecorder的录制格式

移动端的MediaRecorder可能不直接支持audio/wav,需要手动指定兼容的MIME类型。修改MediaRecorder初始化代码:

// 初始化时指定PCM编码的WAV格式
const mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/wav; codecs=1' });

如果当前浏览器不支持该MIME类型,可以降级录制为audio/webm,并在服务端将其转码为LINEAR16格式,或者直接在Speech API中指定encoding=speech.RecognitionConfig.AudioEncoding.OGG_OPUS。

3. 验证前端Blob的实际格式

在上传前添加校验,确认生成的Blob是标准WAV格式:

console.log('Blob type:', blob.type);
// 检查WAV文件头部(标准WAV以RIFF开头,对应字节[82,73,70,70])
blob.arrayBuffer().then(buf => {
  const header = new Uint8Array(buf.slice(0, 4));
  console.log('WAV header bytes:', header);
});

若头部不符合标准,说明MediaRecorder未生成有效WAV,需调整录制配置或进行前端转码。

4. 服务端添加音频格式校验

处理音频前先校验是否为有效WAV文件,避免无效数据传入Speech API:

import wave
from io import BytesIO
from flask import jsonify  # 假设用Flask框架,根据实际框架调整

try:
    with wave.open(BytesIO(audio_bytes), 'rb') as wav_file:
        # 获取音频实际参数并同步到config
        config.sample_rate_hertz = wav_file.getframerate()
        config.audio_channel_count = wav_file.getnchannels()
        config.encoding = speech.RecognitionConfig.AudioEncoding.LINEAR16
except wave.Error:
    return jsonify({'error': '无效的音频格式,请上传标准WAV文件'}), 400

内容的提问来源于stack exchange,提问作者Mike Gvozdev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 16:53:12