在Flutter(Dart)中使用Google Cloud实时语音转文本API实现持续识别
Flutter 集成 Google Cloud 实时语音转文本实现方案
前置准备
- 启用Google Cloud Speech-to-Text API,创建并下载服务账号密钥(JSON格式)
- 在Flutter项目中添加依赖:
dependencies: flutter: sdk: flutter http: ^1.1.0 flutter_sound: ^9.2.13 # 用于捕获麦克风音频流 permission_handler: ^10.2.0 # 处理麦克风权限 - 配置平台权限:
- iOS:在
Info.plist中添加NSMicrophoneUsageDescription字段,说明麦克风使用用途 - Android:在
AndroidManifest.xml中添加<uses-permission android:name="android.permission.RECORD_AUDIO"/>权限声明
- iOS:在
核心代码实现
1. 音频捕获与分片工具类
import 'dart:async'; import 'dart:typed_data'; import 'package:flutter_sound/flutter_sound.dart'; import 'package:permission_handler/permission_handler.dart'; class AudioCaptureService { final FlutterSoundRecorder _recorder = FlutterSoundRecorder(); final StreamController<Uint8List> _audioChunkController = StreamController<Uint8List>.broadcast(); Stream<Uint8List> get audioChunkStream => _audioChunkController.stream; // 初始化录音器,匹配Google Cloud要求的音频格式 Future<void> init() async { var status = await Permission.microphone.request(); if (status != PermissionStatus.granted) throw Exception("麦克风权限未授权"); await _recorder.openRecorder(); // 每2秒生成一个音频分片 await _recorder.setSubscriptionDuration(const Duration(milliseconds: 2000)); await _recorder.startRecorder( toStream: true, codec: Codec.pcm16WAV, // LINEAR16编码,符合Google Cloud要求 sampleRate: 16000, numChannels: 1, ); _recorder.onProgress.listen((event) { if (event.soundBytes != null && event.soundBytes!.isNotEmpty) { _audioChunkController.add(event.soundBytes!); } }); } Future<void> stop() async { await _recorder.stopRecorder(); await _recorder.closeRecorder(); _audioChunkController.close(); } }
2. Google Cloud API调用类
import 'dart:convert'; import 'dart:typed_data'; import 'package:http/http.dart' as http; class GoogleSpeechApi { final String _apiKey; final String _endpoint = "https://speech.googleapis.com/v1p1beta1/speech:recognize"; GoogleSpeechApi(this._apiKey); Future<String> recognizeAudio(Uint8List audioBytes) async { final requestBody = jsonEncode({ "config": { "encoding": "LINEAR16", "sampleRateHertz": 16000, "languageCode": "zh-CN", // 根据需求修改语言代码,如"en-US" "enableAutomaticPunctuation": true }, "audio": { "content": base64Encode(audioBytes) } }); final response = await http.post( Uri.parse("$_endpoint?key=$_apiKey"), headers: {"Content-Type": "application/json"}, body: requestBody, ); if (response.statusCode == 200) { final result = jsonDecode(response.body); if (result['results'] != null && result['results'].isNotEmpty) { return result['results'][0]['alternatives'][0]['transcript']; } return ""; } else { throw Exception("API调用失败: ${response.body}"); } } }
3. 页面UI与逻辑整合
import 'package:flutter/material.dart'; class SpeechRecognitionPage extends StatefulWidget { const SpeechRecognitionPage({super.key}); @override State<SpeechRecognitionPage> createState() => _SpeechRecognitionPageState(); } class _SpeechRecognitionPageState extends State<SpeechRecognitionPage> { final AudioCaptureService _audioService = AudioCaptureService(); // 建议通过安全方式注入密钥,不要硬编码 final GoogleSpeechApi _speechApi = GoogleSpeechApi("YOUR_API_KEY"); String _transcript = ""; bool _isRecording = false; StreamSubscription? _audioChunkSubscription; void _startRecording() async { try { await _audioService.init(); setState(() => _isRecording = true); _audioChunkSubscription = _audioService.audioChunkStream.listen((chunk) async { try { String text = await _speechApi.recognizeAudio(chunk); setState(() => _transcript += "$text "); } catch (e) { debugPrint("识别失败: $e"); } }); } catch (e) { debugPrint("启动录音失败: $e"); } } void _stopRecording() async { await _audioService.stop(); _audioChunkSubscription?.cancel(); setState(() => _isRecording = false); } @override void dispose() { if (_isRecording) _stopRecording(); super.dispose(); } @override Widget build(BuildContext context) { return Scaffold( appBar: AppBar(title: const Text("实时语音转文本")), body: Padding( padding: const EdgeInsets.all(16.0), child: Column( children: [ Expanded( child: SingleChildScrollView( child: Text(_transcript, style: const TextStyle(fontSize: 16)), ), ), const SizedBox(height: 20), ElevatedButton( onPressed: _isRecording ? _stopRecording : _startRecording, child: Text(_isRecording ? "停止识别" : "开始识别"), ), ], ), ), ); } }
关键注意事项
- 音频格式严格匹配:必须使用LINEAR16编码、16000Hz采样率、单声道,否则API会返回错误或识别结果失真
- 密钥安全:禁止硬编码API密钥或服务账号密钥,建议通过后端代理调用API,或使用Flutter环境变量、密钥管理工具注入
- 网络稳定性:实时识别依赖稳定网络,弱网环境下可添加重试机制或缓存逻辑,避免频繁失败
- API配额限制:Google Cloud API有调用频率和配额限制,若2秒间隔的请求触发限制,可适当调整分片时长或合并请求
- 性能优化:音频处理和API调用需放在异步线程执行,避免阻塞UI;可添加防抖逻辑,减少无效请求
- 平台兼容性:测试时覆盖安卓不同版本和iOS设备,确保音频捕获的一致性,部分安卓机型可能需要额外配置音频权限
内容的提问来源于stack exchange,提问作者Aesha
相关产品推荐
相关产品推荐

