React Native Expo录制符合Google Speech-to-Text标准的WAV音频问题排查
录制的WAV音频无法被Google Speech-to-Text API转写,返回空响应
我实现了录音启停功能,本地播放正常,但将录制的WAV音频上传至Google Speech-to-Text API进行转写时返回空响应;而上传网上下载的WAV文件则可正常转写。
React Native 录音与上传代码
const startRecording = async () => { try { console.log("Requesting permissions.."); const { status } = await Audio.requestPermissionsAsync(); if (status !== "granted") { alert("Sorry, we need audio recording permissions to make this work!"); return; } console.log("Starting recording.."); await Audio.setAudioModeAsync({ allowsRecordingIOS: true, playsInSilentModeIOS: true, }); const { recording } = await Audio.Recording.createAsync({ android: { extension: ".wav", outputFormat: Audio.RECORDING_OPTION_ANDROID_OUTPUT_FORMAT_DEFAULT, audioEncoder: Audio.RECORDING_OPTION_ANDROID_AUDIO_ENCODER_DEFAULT, sampleRate: 16000, numberOfChannels: 1, bitRate: 128000, }, ios: { extension: ".wav", audioQuality: Audio.RECORDING_OPTION_IOS_AUDIO_QUALITY_HIGH, sampleRate: 16000, numberOfChannels: 1, bitRate: 128000, linearPCMBitDepth: 16, linearPCMIsBigEndian: false, linearPCMIsFloat: false, }, }); setRecording(recording); } catch (err) { console.error("Failed to start recording", err); } }; const stopRecording = async () => { console.log("Stopping recording.."); if (recording) { await recording.stopAndUnloadAsync(); const uri = recording.getURI(); setRecording(null); uploadAudio(uri); } }; function getOtherUser(data, username) { if (data.receiver.username !== username) { return data.receiver; } if (data.sender.username !== username) { return data.sender; } return null; } const uploadAudio = async (uri) => { const formData = new FormData(); formData.append("file", { uri, name: "recording.wav", type: "audio/wav", }); formData.append("pair_id", details.id); formData.append( "delivered", Object.values(members).some( (entry) => entry.name === receipient?.username, ) ? true : false, ); try { const url = `${API_URL}/api/record_view/`; const response = await fetch(url, { method: "POST", headers: { "Content-Type": "multipart/form-data", Authorization: "Bearer " + token, user: username, }, body: formData, }); const data = await response.json(); if (response.status === 200) { console.log("File successfully uploaded"); } } catch (error) { console.error("Error uploading audio file:", error); } };
Django 后端转写视图代码
import os from django.http import JsonResponse from django.views.decorators.csrf import csrf_exempt from google.cloud import speech_v2 from google.cloud.speech_v2.types import cloud_speech from google.oauth2 import service_account def transcribe_model_selection_v2(project_id: str, model: str, audio_path: str) -> cloud_speech.RecognizeResponse: """Transcribe an audio file.""" # Instantiates a client with credentials credentials = service_account.Credentials.from_service_account_file('service-account-file.json') client = speech_v2.SpeechClient(credentials=credentials) with open(audio_path, "rb") as f: content = f.read() config = cloud_speech.RecognitionConfig( auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(), language_codes=["en-US"], model=model, ) request = cloud_speech.RecognizeRequest( recognizer=f"projects/{project_id}/locations/global/recognizers/_", config=config, content=content, ) response = client.recognize(request=request) return response @csrf_exempt def transcribe_audio(request): if request.method == 'POST' and request.FILES.get('audio'): audio_file = request.FILES['audio'] audio_path = f'/tmp/{audio_file.name}' with open(audio_path, 'wb+') as destination: for chunk in audio_file.chunks(): destination.write(chunk) try: project_id = 'project-id' model = 'latest_long' response = transcribe_model_selection_v2(project_id, model, audio_path) os.remove(audio_path) transcription = '' for result in response.results: transcription += result.alternatives[0].transcript return JsonResponse({'transcription': transcription}) except Exception as e: os.remove(audio_path) return JsonResponse({'error': str(e)}, status=500) return JsonResponse({'error': 'Invalid request'}, status=400)
问题排查与解决方案
1. 修正Android端录音编码格式
Google Speech-to-Text仅支持无压缩的PCM编码WAV,你当前Android端使用的默认编码可能是压缩格式(仅修改了文件后缀)。修改录音配置:
android: { extension: ".wav", // 强制使用PCM 16位编码 outputFormat: Audio.RECORDING_OPTION_ANDROID_OUTPUT_FORMAT_PCM_16BIT, audioEncoder: Audio.RECORDING_OPTION_ANDROID_AUDIO_ENCODER_PCM_16BIT, sampleRate: 16000, numberOfChannels: 1, bitRate: 128000, },
2. 修复FormData上传的Content-Type问题
React Native中手动设置Content-Type: multipart/form-data会丢失自动生成的boundary参数,导致后端无法正确解析文件。移除该设置:
headers: { Authorization: "Bearer " + token, user: username, // 移除手动设置的Content-Type },
3. 验证音频文件完整性
在后端临时保留上传的音频文件,下载后用Audacity等工具检查:
- 是否为PCM编码的WAV格式
- 采样率、声道数是否为16000Hz、单声道
- 文件是否包含有效音频内容(非空或无声)
4. 显式指定Google API的音频配置
替换自动解码配置为显式参数,帮助API精准识别:
config = cloud_speech.RecognitionConfig( encoding=cloud_speech.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_codes=["en-US"], model=model, )
内容的提问来源于stack exchange,提问作者Ajao Malik
相关产品推荐
相关产品推荐

