You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Expo File System读取音频二进制异常致Azure语音转文本错误

问题解决:Azure Speech-To-Text 流式识别返回大量无意义逗号

问题根源分析

你的代码存在三个核心问题,导致Azure识别出无意义逗号:

  1. Android录音格式不兼容:Android端用了MPEG_4输出格式+AAC编码器,生成的是MP4容器的音频(而非真正的WAV/PCM),Azure Speech SDK流式识别仅支持无压缩的PCM格式。
  2. 重复推送音频数据:每隔500ms读取整个录音文件并推送到流,导致每次都重复发送之前的音频内容,音频叠加后被识别为噪音/无意义符号。
  3. 未处理WAV文件头:iOS端生成的LINEARPCM是带44字节WAV头的完整文件,直接推送会让SDK把文件头当成音频数据解析,引发识别错误。

修复后的完整代码

import { useEffect, useRef, useState } from "react";
import { Audio } from "expo-av";
import * as FileSystem from 'expo-file-system';
import { EncodingType } from 'expo-file-system';
import { Platform } from "react-native";
import {
  AudioConfig,
  AudioStreamFormat,
  PushAudioInputStream,
  ResultReason,
  SpeechConfig,
  SpeechRecognizer
} from "microsoft-cognitiveservices-speech-sdk";
import {
  AndroidAudioEncoder,
  AndroidOutputFormat,
  IOSAudioQuality,
  IOSOutputFormat
} from "expo-av/build/Audio/RecordingConstants";

// 统一为Azure兼容的PCM录音格式
const RecordingOptions: Audio.RecordingOptions = {
  isMeteringEnabled: true,
  android: {
    extension: '.raw', // 用raw后缀明确是PCM原始数据
    outputFormat: AndroidOutputFormat.RAW,
    audioEncoder: AndroidAudioEncoder.PCM_16BIT,
    sampleRate: 16000,
    numberOfChannels: 1,
    bitRate: 256000, // 16bit单声道16kHz对应的比特率:16000*1*16=256000
  },
  ios: {
    extension: '.wav',
    outputFormat: IOSOutputFormat.LINEARPCM,
    audioQuality: IOSAudioQuality.HIGH,
    sampleRate: 16000,
    numberOfChannels: 1,
    bitRate: 256000,
    linearPCMBitDepth: 16,
    linearPCMIsBigEndian: false,
    linearPCMIsFloat: false,
  },
  web: {
    mimeType: 'audio/wav; codecs=pcm',
    bitsPerSecond: 256000,
  },
};

export const useAzureSpeechStream = (key: string, region: string, language: "en-US" | "ro-RO") => {
  const [recording, setRecording] = useState<Audio.Recording | null>(null);
  const [isRecording, setIsRecording] = useState<boolean>(false);
  const [transcript, setTranscript] = useState<string>("");
  
  const intervalRef = useRef<NodeJS.Timeout | null>(null);
  const lastProcessedPosition = useRef(0);
  const stream = useRef(PushAudioInputStream.create());
  const recognizer = useRef<SpeechRecognizer | null>(null);

  useEffect(() => {
    // 每次参数变化时重新初始化识别器
    const speechConfig = SpeechConfig.fromSubscription(key, region);
    speechConfig.speechRecognitionLanguage = language;
    
    // 明确指定输入音频格式:16kHz、16bit、单声道PCM
    const audioFormat = AudioStreamFormat.getWaveFormatPCM(16000, 16, 1);
    const audioConfig = AudioConfig.fromStreamInput(stream.current, audioFormat);
    
    const newRecognizer = new SpeechRecognizer(speechConfig, audioConfig);
    
    // 实时识别结果(中间态)
    newRecognizer.recognizing = (_, e) => {
      if (e.result.reason === ResultReason.RecognizingSpeech) {
        console.log(`Recognizing: ${e.result.text}`);
        setTranscript(e.result.text); // 实时更新,避免累加重复内容
      }
    };
    
    // 最终识别结果(确认态)
    newRecognizer.recognized = (_, e) => {
      if (e.result.reason === ResultReason.RecognizedSpeech) {
        console.log(`Recognized: ${e.result.text}`);
        setTranscript(prev => prev + " " + e.result.text);
      }
    };

    newRecognizer.startContinuousRecognitionAsync();
    recognizer.current = newRecognizer;

    return () => {
      newRecognizer.stopContinuousRecognitionAsync();
      newRecognizer.close();
      stream.current.close();
    };
  }, [key, region, language]);
  
  const requestAudioPermissions = async () => {
    const response = await Audio.requestPermissionsAsync();
    return response.status === 'granted';
  };
  
  const startRecording = async () => {
    try {
      if (recording) {
        await recording.stopAndUnloadAsync();
        setRecording(null);
      }
      
      const hasPermission = await requestAudioPermissions();
      if (!hasPermission) {
        throw new Error('Permission to access microphone is required!');
      }
      
      await Audio.setAudioModeAsync({
        allowsRecordingIOS: true,
        playsInSilentModeIOS: true,
        staysActiveInBackground: true,
        shouldDuckAndroid: true,
      });
      
      const newRecording = new Audio.Recording();
      await newRecording.prepareToRecordAsync(RecordingOptions);
      await newRecording.startAsync();
      setRecording(newRecording);
      setIsRecording(true);
      console.log('Recording started');
      
      // 重置已读取位置
      lastProcessedPosition.current = 0;
      
      intervalRef.current = setInterval(async () => {
        const uri = newRecording.getURI();
        if (!uri) return;
        
        const fileInfo = await FileSystem.getInfoAsync(uri, {size: true});
        if (!fileInfo.exists || !fileInfo.size) return;
        
        const currentSize = fileInfo.size;
        const bytesToRead = currentSize - lastProcessedPosition.current;
        
        if (bytesToRead <= 0) return;
        
        // 仅读取新增的音频数据
        const fileString = await FileSystem.readAsStringAsync(uri, {
          encoding: EncodingType.Base64,
          position: lastProcessedPosition.current,
          length: bytesToRead,
        });
        
        let fileBuffer = Buffer.from(fileString, 'base64');
        
        // iOS端跳过WAV文件头(前44字节),Android是RAW PCM无需处理
        if (Platform.OS === 'ios' && lastProcessedPosition.current === 0) {
          fileBuffer = fileBuffer.slice(44);
        }
        
        // 推送纯PCM数据到Azure流
        stream.current.write(fileBuffer);
        
        // 更新已读取位置
        lastProcessedPosition.current = currentSize;
      }, 500);
      
    } catch (error) {
      console.error('Failed to start recording:', error);
      setIsRecording(false);
      throw error;
    }
  };
  
  const stopRecording = async () => {
    if (!recording) return;
    
    try {
      await recording.stopAndUnloadAsync();
      setRecording(null);
      setIsRecording(false);
      
      if (intervalRef.current) {
        clearInterval(intervalRef.current);
        intervalRef.current = null;
      }
      
      stream.current.close();
      recognizer.current?.stopContinuousRecognitionAsync();
    } catch (error) {
      console.error("Error stopping recording:", error);
      throw error;
    }
  };
  
  return {
    isRecording,
    transcript,
    startRecording,
    stopRecording,
  };
};

关键修复点说明

  1. 统一音频格式:Android端改为RAW输出格式+PCM_16BIT编码器,和iOS的LINEARPCM保持一致,确保输出Azure兼容的无压缩PCM数据。
  2. 增量读取音频:记录上次读取的文件位置,每次仅推送新增的音频内容,避免重复发送导致的音频叠加。
  3. 处理WAV文件头:iOS端生成的WAV文件包含44字节的文件头,第一次读取时跳过,只推送纯PCM音频数据。
  4. 明确指定音频格式:创建AudioStreamFormat告诉Azure SDK输入音频的规格(采样率、位深、声道数),避免自动解析错误。
  5. 修复识别器依赖:将识别器初始化逻辑放入useEffect并添加参数依赖,确保参数变化时重新创建识别器。

内容的提问来源于stack exchange,提问作者Bogdan Costin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 06:44:54