You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

React Native Expo录制符合Google Speech-to-Text标准的WAV音频问题排查

录制的WAV音频无法被Google Speech-to-Text API转写,返回空响应

我实现了录音启停功能,本地播放正常,但将录制的WAV音频上传至Google Speech-to-Text API进行转写时返回空响应;而上传网上下载的WAV文件则可正常转写。


React Native 录音与上传代码

const startRecording = async () => {
    try {
      console.log("Requesting permissions..");
      const { status } = await Audio.requestPermissionsAsync();
      if (status !== "granted") {
        alert("Sorry, we need audio recording permissions to make this work!");
        return;
      }

      console.log("Starting recording..");
      await Audio.setAudioModeAsync({
        allowsRecordingIOS: true,
        playsInSilentModeIOS: true,
      });

      const { recording } = await Audio.Recording.createAsync({
        android: {
          extension: ".wav",
          outputFormat: Audio.RECORDING_OPTION_ANDROID_OUTPUT_FORMAT_DEFAULT,
          audioEncoder: Audio.RECORDING_OPTION_ANDROID_AUDIO_ENCODER_DEFAULT,
          sampleRate: 16000,
          numberOfChannels: 1,
          bitRate: 128000,
        },
        ios: {
          extension: ".wav",
          audioQuality: Audio.RECORDING_OPTION_IOS_AUDIO_QUALITY_HIGH,
          sampleRate: 16000,
          numberOfChannels: 1,
          bitRate: 128000,
          linearPCMBitDepth: 16,
          linearPCMIsBigEndian: false,
          linearPCMIsFloat: false,
        },
      });
      setRecording(recording);
    } catch (err) {
      console.error("Failed to start recording", err);
    }
  };

  const stopRecording = async () => {
    console.log("Stopping recording..");
    if (recording) {
      await recording.stopAndUnloadAsync();
      const uri = recording.getURI();
      setRecording(null);
      uploadAudio(uri);
    }
  };

  function getOtherUser(data, username) {
    if (data.receiver.username !== username) {
      return data.receiver;
    }
    if (data.sender.username !== username) {
      return data.sender;
    }
    return null;
  }

  const uploadAudio = async (uri) => {
    const formData = new FormData();
    formData.append("file", {
      uri,
      name: "recording.wav",
      type: "audio/wav",
    });
    formData.append("pair_id", details.id);
    formData.append(
      "delivered",
      Object.values(members).some(
        (entry) => entry.name === receipient?.username,
      )
        ? true
        : false,
    );

    try {
      const url = `${API_URL}/api/record_view/`;
      const response = await fetch(url, {
        method: "POST",
        headers: {
          "Content-Type": "multipart/form-data",
          Authorization: "Bearer " + token,
          user: username,
        },
        body: formData,
      });
      const data = await response.json();
      if (response.status === 200) {
        console.log("File successfully uploaded");
      }
    } catch (error) {
      console.error("Error uploading audio file:", error);
    }
  };

Django 后端转写视图代码

import os
from django.http import JsonResponse
from django.views.decorators.csrf import csrf_exempt
from google.cloud import speech_v2
from google.cloud.speech_v2.types import cloud_speech
from google.oauth2 import service_account

def transcribe_model_selection_v2(project_id: str, model: str, audio_path: str) -> cloud_speech.RecognizeResponse:
    """Transcribe an audio file."""
    # Instantiates a client with credentials
    credentials = service_account.Credentials.from_service_account_file('service-account-file.json')
    client = speech_v2.SpeechClient(credentials=credentials)

    with open(audio_path, "rb") as f:
        content = f.read()

    config = cloud_speech.RecognitionConfig(
        auto_decoding_config=cloud_speech.AutoDetectDecodingConfig(),
        language_codes=["en-US"],
        model=model,
    )

    request = cloud_speech.RecognizeRequest(
        recognizer=f"projects/{project_id}/locations/global/recognizers/_",
        config=config,
        content=content,
    )

    response = client.recognize(request=request)

    return response

@csrf_exempt
def transcribe_audio(request):
    if request.method == 'POST' and request.FILES.get('audio'):
        audio_file = request.FILES['audio']
        audio_path = f'/tmp/{audio_file.name}'
        
        with open(audio_path, 'wb+') as destination:
            for chunk in audio_file.chunks():
                destination.write(chunk)

        try:
            project_id = 'project-id'
            model = 'latest_long' 
            response = transcribe_model_selection_v2(project_id, model, audio_path)

            os.remove(audio_path)

            transcription = ''
            for result in response.results:
                transcription += result.alternatives[0].transcript

            return JsonResponse({'transcription': transcription})

        except Exception as e:
            os.remove(audio_path)
            return JsonResponse({'error': str(e)}, status=500)

    return JsonResponse({'error': 'Invalid request'}, status=400)

问题排查与解决方案

1. 修正Android端录音编码格式

Google Speech-to-Text仅支持无压缩的PCM编码WAV,你当前Android端使用的默认编码可能是压缩格式(仅修改了文件后缀)。修改录音配置:

android: {
  extension: ".wav",
  // 强制使用PCM 16位编码
  outputFormat: Audio.RECORDING_OPTION_ANDROID_OUTPUT_FORMAT_PCM_16BIT,
  audioEncoder: Audio.RECORDING_OPTION_ANDROID_AUDIO_ENCODER_PCM_16BIT,
  sampleRate: 16000,
  numberOfChannels: 1,
  bitRate: 128000,
},

2. 修复FormData上传的Content-Type问题

React Native中手动设置Content-Type: multipart/form-data会丢失自动生成的boundary参数,导致后端无法正确解析文件。移除该设置:

headers: {
  Authorization: "Bearer " + token,
  user: username,
  // 移除手动设置的Content-Type
},

3. 验证音频文件完整性

在后端临时保留上传的音频文件,下载后用Audacity等工具检查:

  • 是否为PCM编码的WAV格式
  • 采样率、声道数是否为16000Hz、单声道
  • 文件是否包含有效音频内容(非空或无声)

4. 显式指定Google API的音频配置

替换自动解码配置为显式参数,帮助API精准识别:

config = cloud_speech.RecognitionConfig(
    encoding=cloud_speech.RecognitionConfig.AudioEncoding.LINEAR16,
    sample_rate_hertz=16000,
    language_codes=["en-US"],
    model=model,
)

内容的提问来源于stack exchange,提问作者Ajao Malik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 20:43:16