You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Expo React Native无需后端实现Azure语音转文本及问题咨询

问题解答:Expo + React Native 语音转文本实现及Azure SDK报错处理

一、Azure语音SDK能否直接使用Expo录制的本地WAV文件?

可以直接使用,但要注意文件路径的适配——microsoft-cognitiveservices-speech-sdk的AudioConfig.fromWavFileInput方法在React Native/Expo环境下,无法直接识别Expo返回的file:///格式URI,它需要的是文件的二进制流数据,这是你触发TypeError: Cannot read property 'slice' of undefined错误的核心原因。

二、解决当前报错的具体方案

你的代码存在两个关键问题,针对性修复如下:

1. 修复stopRecording的状态更新顺序

你在停止录制时先执行了setRecording(undefined),导致后续getTranscription调用时recording已经是undefined,自然无法调用getURI()。调整顺序:

async function stopRecording() {
  console.log('Stopping recording..');
  // 先获取录制文件URI,再更新状态
  const uri = recording.getURI();
  await recording.stopAndUnloadAsync();
  await Audio.setAudioModeAsync({ allowsRecordingIOS: false });
  setRecording(undefined); // 最后更新状态
  console.log('Recording stopped and stored at', uri);

  try {
    await getTranscription(uri); // 将URI传入转录函数
  } catch (error) {
    console.log(error);
  }
}

2. 修改getTranscription,读取文件二进制数据

先安装expo-file-system用于读取本地文件:

npx expo install expo-file-system

然后修改转录函数,将URI转换为SDK支持的二进制格式:

import * as FileSystem from 'expo-file-system';

async function getTranscription(uri) {
  try {
    // 读取文件为Base64,再转换为ArrayBuffer
    const base64Content = await FileSystem.readAsStringAsync(uri, {
      encoding: FileSystem.EncodingType.Base64,
    });
    const arrayBuffer = Buffer.from(base64Content, 'base64');
    
    // 创建符合要求的AudioConfig
    const audioConfig = AudioConfig.fromWavFileInput(arrayBuffer);
    const speechConfig = SpeechConfig.fromSubscription('你的Azure语音密钥', 'eastus');
    speechConfig.speechRecognitionLanguage = "en-US";
    
    const recognizer = new SpeechRecognizer(speechConfig, audioConfig);
    
    // 执行语音识别
    recognizer.recognizeOnceAsync(result => {
      switch (result.reason) {
        case ResultReason.RecognizedSpeech:
          console.log(`识别结果: ${result.text}`);
          setMessages(prev => [...prev, result.text]); // 更新消息列表
          break;
        case ResultReason.NoMatch:
          console.log("无法识别语音内容");
          break;
        case ResultReason.Canceled:
          const cancellation = CancellationDetails.fromResult(result);
          console.log(`识别取消原因: ${cancellation.reason}`);
          if (cancellation.reason === CancellationReason.Error) {
            console.log(`错误码: ${cancellation.errorCode}`);
            console.log(`错误详情: ${cancellation.errorDetails}`);
          }
          break;
      }
      recognizer.close();
    });
  } catch (err) {
    console.error('处理音频文件失败:', err);
  }
}

三、兼容Expo/React Native的其他语音转文本包

  • Expo Speech: 官方原生集成的API,无需额外配置,支持多语言实时识别,调用Speech.startRecognitionAsync即可快速实现功能,适合Demo或轻量应用。
  • react-native-voice: 社区维护的原生模块,支持部分平台离线识别,功能更灵活,但Expo环境下需要使用expo-dev-client或 eject 项目才能兼容。
  • Google Cloud Speech-to-Text: 通过HTTP API调用,需自行处理音频文件上传,适合需要Google生态集成的场景,无需依赖原生SDK。

内容的提问来源于stack exchange,提问作者Isuru Akalanka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 18:57:46