Google Cloud Speech API桌面可用但iPhone Chrome移动端报错求助
问题背景
Web服务接收用户音频并通过Google Cloud Speech API转写文本,桌面端(本地/服务器部署)运行正常,但在iPhone的Chrome浏览器中调用时返回:google.api_core.exceptions.InvalidArgument: 400 RecognitionAudio not set.
移动端请求日志额外显示:details = "Invalid recognition 'config': bad encoding.."
已确认服务端收到的音频数据非空,但移动端录制的音频存在编码适配问题。
前端录制代码(JavaScript)
// Store the audio data in chunks as it is recorded mediaRecorder.addEventListener("dataavailable", (event) => { chunks.push(event.data); }); // When recording stops, create a blob from the chunks and send it to the server mediaRecorder.addEventListener("stop", () => { const blob = new Blob(chunks, { type: "audio/wav" }); const url = URL.createObjectURL(blob); audio.src = url; uploadAudio(blob); // Send the audio data to the server chunks = []; // Clear the chunks array });
服务端处理代码(Python)
audio_bytes = request.get_data() # Instantiate a client client = speech.SpeechClient() # Settings for Speech v1: # Initialize request arguments config = speech.RecognitionConfig() config.language_code = LANGUAGE config.model = "latest_long" config.enable_automatic_punctuation = True audio = speech.RecognitionAudio() audio.content = audio_bytes speech_request = speech.RecognizeRequest( config=config, audio=audio, ) # Make the request response = client.recognize(request=speech_request)
解决方案
1. 强制指定Speech API的音频编码参数
尽管文档说明WAV无需设置编码,但移动端浏览器生成的音频格式可能和桌面端存在差异(比如iPhone Chrome默认用AAC编码而非PCM)。在服务端的RecognitionConfig中明确添加编码配置:
config.audio_channel_count = 1 # 适配移动端常见的单声道录制 config.sample_rate_hertz = 16000 # Speech API推荐的16kHz采样率 config.encoding = speech.RecognitionConfig.AudioEncoding.LINEAR16 # WAV对应的PCM编码
2. 修正前端MediaRecorder的录制格式
移动端的MediaRecorder可能不直接支持audio/wav,需要手动指定兼容的MIME类型。修改MediaRecorder初始化代码:
// 初始化时指定PCM编码的WAV格式 const mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/wav; codecs=1' });
如果当前浏览器不支持该MIME类型,可以降级录制为audio/webm,并在服务端将其转码为LINEAR16格式,或者直接在Speech API中指定encoding=speech.RecognitionConfig.AudioEncoding.OGG_OPUS。
3. 验证前端Blob的实际格式
在上传前添加校验,确认生成的Blob是标准WAV格式:
console.log('Blob type:', blob.type); // 检查WAV文件头部(标准WAV以RIFF开头,对应字节[82,73,70,70]) blob.arrayBuffer().then(buf => { const header = new Uint8Array(buf.slice(0, 4)); console.log('WAV header bytes:', header); });
若头部不符合标准,说明MediaRecorder未生成有效WAV,需调整录制配置或进行前端转码。
4. 服务端添加音频格式校验
处理音频前先校验是否为有效WAV文件,避免无效数据传入Speech API:
import wave from io import BytesIO from flask import jsonify # 假设用Flask框架,根据实际框架调整 try: with wave.open(BytesIO(audio_bytes), 'rb') as wav_file: # 获取音频实际参数并同步到config config.sample_rate_hertz = wav_file.getframerate() config.audio_channel_count = wav_file.getnchannels() config.encoding = speech.RecognitionConfig.AudioEncoding.LINEAR16 except wave.Error: return jsonify({'error': '无效的音频格式,请上传标准WAV文件'}), 400
内容的提问来源于stack exchange,提问作者Mike Gvozdev

