React Native本地离线提取音频文本:求解决方案及代码示例
React Native 本地离线音频转文字实现方案
完全可以在React Native中实现离线本地音频转文字,不需要联网。主要有两种可行思路,分别基于系统原生API和开源本地语音模型,以下是具体方案和代码示例:
方案一:利用系统原生离线语音识别API
iOS和Android系统本身提供了离线语音识别能力,只需提前下载对应语言的离线识别包,通过RN桥接库即可调用。
依赖库:react-native-voice
这个库封装了iOS的SFSpeechRecognizer和Android的SpeechRecognizer,支持离线模式。
实现步骤与代码
- 安装依赖:
npm install @react-native-voice/voice --save # 或 yarn add @react-native-voice/voice
- 权限配置:
- iOS:在
Info.plist中添加权限声明:<key>NSMicrophoneUsageDescription</key> <string>需要使用麦克风进行语音识别</string> <key>NSSpeechRecognitionUsageDescription</key> <string>需要使用语音识别功能</string> - Android:在
AndroidManifest.xml中添加权限:
同时确保设备已下载对应语言的离线语音包(iOS在「设置-通用-键盘-语音」中下载;Android在系统语言设置中下载离线识别包)。<uses-permission android:name="android.permission.RECORD_AUDIO" />
- 核心代码示例:
import React, { Component } from 'react'; import { View, Text, Button } from 'react-native'; import Voice from '@react-native-voice/voice'; class OfflineVoiceToText extends Component { state = { transcription: '', isListening: false, }; componentDidMount() { // 绑定识别事件 Voice.onSpeechStart = this.handleSpeechStart; Voice.onSpeechEnd = this.handleSpeechEnd; Voice.onSpeechResults = this.handleSpeechResults; // 请求权限并开启离线模式 Voice.requestPermissions().catch(err => console.error(err)); Voice.setOfflineMode(true); } componentWillUnmount() { // 清理资源 Voice.destroy().then(Voice.removeAllListeners); } handleSpeechStart = () => this.setState({ isListening: true }); handleSpeechEnd = () => this.setState({ isListening: false }); handleSpeechResults = (event) => { this.setState({ transcription: event.value[0], }); }; startRecognition = async () => { try { await Voice.start('zh-CN'); // 指定中文离线识别,可替换为其他语言码 } catch (err) { console.error('启动识别失败:', err); } }; stopRecognition = async () => { try { await Voice.stop(); } catch (err) { console.error('停止识别失败:', err); } }; render() { return ( <View style={{ padding: 20 }}> <Text style={{ fontSize: 16, marginBottom: 20 }}>转录结果:{this.state.transcription}</Text> <Button title={this.state.isListening ? '停止识别' : '开始识别'} onPress={this.state.isListening ? this.stopRecognition : this.startRecognition} /> </View> ); } } export default OfflineVoiceToText;
注意事项
- 部分Android厂商对系统语音识别的支持程度不同,离线语言包的可用性可能有差异。
- iOS需要用户手动下载离线语音包,无法通过代码自动触发下载。
方案二:集成Whisper.cpp本地模型(完全离线跨平台)
如果需要摆脱系统依赖,完全自主实现离线转文字,可基于OpenAI开源的Whisper模型,通过react-native-whisper封装库在RN中运行本地模型。
依赖库:react-native-whisper
该库封装了Whisper的C++实现,支持本地加载模型,无需联网,兼容多语言。
实现步骤与代码
- 安装依赖:
npm install react-native-whisper --save
下载模型:
从Whisper官方仓库下载预训练模型(推荐ggml-base.bin,约1GB,支持多语言),将模型文件放在项目的assets目录下。配置模型路径:
- iOS:在Xcode中把模型文件添加到「Copy Bundle Resources」中。
- Android:将模型文件放在
android/app/src/main/assets目录。
- 核心代码示例:
import React, { useState } from 'react'; import { View, Text, Button, TextInput } from 'react-native'; import Whisper from 'react-native-whisper'; const WhisperTranscriber = () => { const [audioPath, setAudioPath] = useState(''); const [transcription, setTranscription] = useState(''); const transcribeLocalAudio = async () => { if (!audioPath) return; try { // 预加载模型(建议在APP启动时提前加载) await Whisper.loadModel(require('./assets/ggml-base.bin')); // 转录本地音频文件(支持mp3、wav等格式) const result = await Whisper.transcribe(audioPath); setTranscription(result.text); } catch (err) { console.error('转录失败:', err); setTranscription('转录出错,请检查音频文件'); } }; return ( <View style={{ padding: 20 }}> <TextInput placeholder="输入本地音频文件路径" value={audioPath} onChangeText={setAudioPath} style={{ borderBottomWidth: 1, marginBottom: 20, padding: 10 }} /> <Button title="开始转录" onPress={transcribeLocalAudio} /> <Text style={{ marginTop: 20, fontSize: 16 }}>转录结果:{transcription}</Text> </View> ); }; export default WhisperTranscriber;
注意事项
- 模型文件较大,会增加APP安装包体积,可根据需求选择更小的模型(如
ggml-tiny.bin,约70MB,但识别精度略低)。 - 首次加载模型需要几秒时间,建议在APP启动时预加载提升体验。
- 部分音频格式可能需要转码为16kHz单声道WAV,可配合
react-native-audio-toolkit等库处理音频。
方案对比
| 方案类型 | 优点 | 缺点 |
|---|---|---|
| 系统原生API | 体积小、识别速度快 | 依赖系统支持、离线包需手动下载 |
| Whisper本地模型 | 完全离线、跨平台兼容好、多语言支持 | APP体积增大、首次加载慢 |
可根据项目需求选择合适的方案。
内容的提问来源于stack exchange,提问作者Ahsen Shiekh
相关产品推荐
相关产品推荐

