You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Flutter(Dart)中使用Google Cloud实时语音转文本API实现持续识别

Flutter 集成 Google Cloud 实时语音转文本实现方案

前置准备

  • 启用Google Cloud Speech-to-Text API,创建并下载服务账号密钥(JSON格式)
  • 在Flutter项目中添加依赖:
    dependencies:
      flutter:
        sdk: flutter
      http: ^1.1.0
      flutter_sound: ^9.2.13 # 用于捕获麦克风音频流
      permission_handler: ^10.2.0 # 处理麦克风权限
    
  • 配置平台权限:
    • iOS:在Info.plist中添加NSMicrophoneUsageDescription字段,说明麦克风使用用途
    • Android:在AndroidManifest.xml中添加<uses-permission android:name="android.permission.RECORD_AUDIO"/>权限声明

核心代码实现

1. 音频捕获与分片工具类

import 'dart:async';
import 'dart:typed_data';
import 'package:flutter_sound/flutter_sound.dart';
import 'package:permission_handler/permission_handler.dart';

class AudioCaptureService {
  final FlutterSoundRecorder _recorder = FlutterSoundRecorder();
  final StreamController<Uint8List> _audioChunkController = StreamController<Uint8List>.broadcast();
  Stream<Uint8List> get audioChunkStream => _audioChunkController.stream;

  // 初始化录音器,匹配Google Cloud要求的音频格式
  Future<void> init() async {
    var status = await Permission.microphone.request();
    if (status != PermissionStatus.granted) throw Exception("麦克风权限未授权");
    
    await _recorder.openRecorder();
    // 每2秒生成一个音频分片
    await _recorder.setSubscriptionDuration(const Duration(milliseconds: 2000));
    await _recorder.startRecorder(
      toStream: true,
      codec: Codec.pcm16WAV, // LINEAR16编码,符合Google Cloud要求
      sampleRate: 16000,
      numChannels: 1,
    );
    
    _recorder.onProgress.listen((event) {
      if (event.soundBytes != null && event.soundBytes!.isNotEmpty) {
        _audioChunkController.add(event.soundBytes!);
      }
    });
  }

  Future<void> stop() async {
    await _recorder.stopRecorder();
    await _recorder.closeRecorder();
    _audioChunkController.close();
  }
}

2. Google Cloud API调用类

import 'dart:convert';
import 'dart:typed_data';
import 'package:http/http.dart' as http;

class GoogleSpeechApi {
  final String _apiKey;
  final String _endpoint = "https://speech.googleapis.com/v1p1beta1/speech:recognize";

  GoogleSpeechApi(this._apiKey);

  Future<String> recognizeAudio(Uint8List audioBytes) async {
    final requestBody = jsonEncode({
      "config": {
        "encoding": "LINEAR16",
        "sampleRateHertz": 16000,
        "languageCode": "zh-CN", // 根据需求修改语言代码,如"en-US"
        "enableAutomaticPunctuation": true
      },
      "audio": {
        "content": base64Encode(audioBytes)
      }
    });

    final response = await http.post(
      Uri.parse("$_endpoint?key=$_apiKey"),
      headers: {"Content-Type": "application/json"},
      body: requestBody,
    );

    if (response.statusCode == 200) {
      final result = jsonDecode(response.body);
      if (result['results'] != null && result['results'].isNotEmpty) {
        return result['results'][0]['alternatives'][0]['transcript'];
      }
      return "";
    } else {
      throw Exception("API调用失败: ${response.body}");
    }
  }
}

3. 页面UI与逻辑整合

import 'package:flutter/material.dart';

class SpeechRecognitionPage extends StatefulWidget {
  const SpeechRecognitionPage({super.key});

  @override
  State<SpeechRecognitionPage> createState() => _SpeechRecognitionPageState();
}

class _SpeechRecognitionPageState extends State<SpeechRecognitionPage> {
  final AudioCaptureService _audioService = AudioCaptureService();
  // 建议通过安全方式注入密钥,不要硬编码
  final GoogleSpeechApi _speechApi = GoogleSpeechApi("YOUR_API_KEY");
  String _transcript = "";
  bool _isRecording = false;
  StreamSubscription? _audioChunkSubscription;

  void _startRecording() async {
    try {
      await _audioService.init();
      setState(() => _isRecording = true);
      _audioChunkSubscription = _audioService.audioChunkStream.listen((chunk) async {
        try {
          String text = await _speechApi.recognizeAudio(chunk);
          setState(() => _transcript += "$text ");
        } catch (e) {
          debugPrint("识别失败: $e");
        }
      });
    } catch (e) {
      debugPrint("启动录音失败: $e");
    }
  }

  void _stopRecording() async {
    await _audioService.stop();
    _audioChunkSubscription?.cancel();
    setState(() => _isRecording = false);
  }

  @override
  void dispose() {
    if (_isRecording) _stopRecording();
    super.dispose();
  }

  @override
  Widget build(BuildContext context) {
    return Scaffold(
      appBar: AppBar(title: const Text("实时语音转文本")),
      body: Padding(
        padding: const EdgeInsets.all(16.0),
        child: Column(
          children: [
            Expanded(
              child: SingleChildScrollView(
                child: Text(_transcript, style: const TextStyle(fontSize: 16)),
              ),
            ),
            const SizedBox(height: 20),
            ElevatedButton(
              onPressed: _isRecording ? _stopRecording : _startRecording,
              child: Text(_isRecording ? "停止识别" : "开始识别"),
            ),
          ],
        ),
      ),
    );
  }
}

关键注意事项

  • 音频格式严格匹配:必须使用LINEAR16编码、16000Hz采样率、单声道,否则API会返回错误或识别结果失真
  • 密钥安全:禁止硬编码API密钥或服务账号密钥,建议通过后端代理调用API,或使用Flutter环境变量、密钥管理工具注入
  • 网络稳定性:实时识别依赖稳定网络,弱网环境下可添加重试机制或缓存逻辑,避免频繁失败
  • API配额限制:Google Cloud API有调用频率和配额限制,若2秒间隔的请求触发限制,可适当调整分片时长或合并请求
  • 性能优化:音频处理和API调用需放在异步线程执行,避免阻塞UI;可添加防抖逻辑,减少无效请求
  • 平台兼容性:测试时覆盖安卓不同版本和iOS设备,确保音频捕获的一致性,部分安卓机型可能需要额外配置音频权限

内容的提问来源于stack exchange,提问作者Aesha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 11:00:32