You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

React Native本地离线提取音频文本:求解决方案及代码示例

React Native 本地离线音频转文字实现方案

完全可以在React Native中实现离线本地音频转文字,不需要联网。主要有两种可行思路,分别基于系统原生API和开源本地语音模型,以下是具体方案和代码示例:

方案一:利用系统原生离线语音识别API

iOS和Android系统本身提供了离线语音识别能力,只需提前下载对应语言的离线识别包,通过RN桥接库即可调用。

依赖库:react-native-voice

这个库封装了iOS的SFSpeechRecognizer和Android的SpeechRecognizer,支持离线模式。

实现步骤与代码

  1. 安装依赖:
npm install @react-native-voice/voice --save
# 或 yarn add @react-native-voice/voice
  1. 权限配置:
  • iOS:在Info.plist中添加权限声明:
    <key>NSMicrophoneUsageDescription</key>
    <string>需要使用麦克风进行语音识别</string>
    <key>NSSpeechRecognitionUsageDescription</key>
    <string>需要使用语音识别功能</string>
    
  • Android:在AndroidManifest.xml中添加权限:
    <uses-permission android:name="android.permission.RECORD_AUDIO" />
    
    同时确保设备已下载对应语言的离线语音包(iOS在「设置-通用-键盘-语音」中下载;Android在系统语言设置中下载离线识别包)。
  1. 核心代码示例:
import React, { Component } from 'react';
import { View, Text, Button } from 'react-native';
import Voice from '@react-native-voice/voice';

class OfflineVoiceToText extends Component {
  state = {
    transcription: '',
    isListening: false,
  };

  componentDidMount() {
    // 绑定识别事件
    Voice.onSpeechStart = this.handleSpeechStart;
    Voice.onSpeechEnd = this.handleSpeechEnd;
    Voice.onSpeechResults = this.handleSpeechResults;
    
    // 请求权限并开启离线模式
    Voice.requestPermissions().catch(err => console.error(err));
    Voice.setOfflineMode(true);
  }

  componentWillUnmount() {
    // 清理资源
    Voice.destroy().then(Voice.removeAllListeners);
  }

  handleSpeechStart = () => this.setState({ isListening: true });
  handleSpeechEnd = () => this.setState({ isListening: false });

  handleSpeechResults = (event) => {
    this.setState({
      transcription: event.value[0],
    });
  };

  startRecognition = async () => {
    try {
      await Voice.start('zh-CN'); // 指定中文离线识别,可替换为其他语言码
    } catch (err) {
      console.error('启动识别失败:', err);
    }
  };

  stopRecognition = async () => {
    try {
      await Voice.stop();
    } catch (err) {
      console.error('停止识别失败:', err);
    }
  };

  render() {
    return (
      <View style={{ padding: 20 }}>
        <Text style={{ fontSize: 16, marginBottom: 20 }}>转录结果:{this.state.transcription}</Text>
        <Button
          title={this.state.isListening ? '停止识别' : '开始识别'}
          onPress={this.state.isListening ? this.stopRecognition : this.startRecognition}
        />
      </View>
    );
  }
}

export default OfflineVoiceToText;

注意事项

  • 部分Android厂商对系统语音识别的支持程度不同,离线语言包的可用性可能有差异。
  • iOS需要用户手动下载离线语音包,无法通过代码自动触发下载。

方案二:集成Whisper.cpp本地模型(完全离线跨平台)

如果需要摆脱系统依赖,完全自主实现离线转文字,可基于OpenAI开源的Whisper模型,通过react-native-whisper封装库在RN中运行本地模型。

依赖库:react-native-whisper

该库封装了Whisper的C++实现,支持本地加载模型,无需联网,兼容多语言。

实现步骤与代码

  1. 安装依赖:
npm install react-native-whisper --save
  1. 下载模型:
    从Whisper官方仓库下载预训练模型(推荐ggml-base.bin,约1GB,支持多语言),将模型文件放在项目的assets目录下。

  2. 配置模型路径:

  • iOS:在Xcode中把模型文件添加到「Copy Bundle Resources」中。
  • Android:将模型文件放在android/app/src/main/assets目录。
  1. 核心代码示例:
import React, { useState } from 'react';
import { View, Text, Button, TextInput } from 'react-native';
import Whisper from 'react-native-whisper';

const WhisperTranscriber = () => {
  const [audioPath, setAudioPath] = useState('');
  const [transcription, setTranscription] = useState('');

  const transcribeLocalAudio = async () => {
    if (!audioPath) return;
    try {
      // 预加载模型(建议在APP启动时提前加载)
      await Whisper.loadModel(require('./assets/ggml-base.bin'));
      // 转录本地音频文件(支持mp3、wav等格式)
      const result = await Whisper.transcribe(audioPath);
      setTranscription(result.text);
    } catch (err) {
      console.error('转录失败:', err);
      setTranscription('转录出错,请检查音频文件');
    }
  };

  return (
    <View style={{ padding: 20 }}>
      <TextInput
        placeholder="输入本地音频文件路径"
        value={audioPath}
        onChangeText={setAudioPath}
        style={{ borderBottomWidth: 1, marginBottom: 20, padding: 10 }}
      />
      <Button title="开始转录" onPress={transcribeLocalAudio} />
      <Text style={{ marginTop: 20, fontSize: 16 }}>转录结果:{transcription}</Text>
    </View>
  );
};

export default WhisperTranscriber;

注意事项

  • 模型文件较大,会增加APP安装包体积,可根据需求选择更小的模型(如ggml-tiny.bin,约70MB,但识别精度略低)。
  • 首次加载模型需要几秒时间,建议在APP启动时预加载提升体验。
  • 部分音频格式可能需要转码为16kHz单声道WAV,可配合react-native-audio-toolkit等库处理音频。

方案对比

方案类型优点缺点
系统原生API体积小、识别速度快依赖系统支持、离线包需手动下载
Whisper本地模型完全离线、跨平台兼容好、多语言支持APP体积增大、首次加载慢

可根据项目需求选择合适的方案。

内容的提问来源于stack exchange,提问作者Ahsen Shiekh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 07:31:43