You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flutter中如何获取音频流以对接Azure认知语音服务实现实时转写?

问题

我正尝试在Flutter中使用Azure认知语音服务实现实时转写功能,需要通过API将音频数据流发送至Azure,但使用record包无法获取音频流。请问是否需要使用其他适配的包?

获取令牌与调用API的代码

String authToken = '';

Future<void> fetchToken() async {
  String subscriptionKey = 'YOUR_SUBSCRIPTION_KEY';
  String region = 'YOUR_REGION';

  final response = await http.post(
      Uri.parse('https://$region.api.cognitive.microsoft.com/sts/v1.0/issueToken'),
      headers: {'Ocp-Apim-Subscription-Key': subscriptionKey}
  );

  if (response.statusCode == 200) {
    authToken = response.body;
  } else {
    throw Exception('Failed to obtain authorization token: ${response.statusCode}');
  }
}


String endpoint = 'https://YOUR_REGION.stt.speech.microsoft.com/speech/recognition/conversation/cognitiveservices/v1';
Uri uri = Uri.parse(endpoint);
Map<String, String> headers = {
  'Content-Type': 'audio/wav',
  'Authorization': 'Bearer $authToken',
};
Map<String, String> params = {
  'language': 'en-US',
  'format': 'simple',
  'profanity': 'masked',
  'recognitionMode': 'conversation',
};
Stream<List<int>> audioStream = getMicrophoneStream();
http.MultipartRequest request = http.MultipartRequest('POST', uri)
  ..headers.addAll(headers)
  ..fields.addAll(params)
  ..files.add(await http.MultipartFile.fromStream('audio', audioStream, contentType: MediaType('audio', 'wav')));

录制音频的代码

Future<void> _start() async {
Directory? directory;
var status = await Permission.storage.status;
try {
  if (Platform.isAndroid) {
    if (status.isGranted) {
      directory = (await getExternalStorageDirectory())!;
      String newPath = "${directory.path}/Recordings";
      directory = Directory(newPath);
    } else {
      await Permission.storage.request();
    }
  } else {
    if (status.isGranted) {
      directory = await getApplicationDocumentsDirectory();
      String newPath = "${directory.path}/Recordings";
      directory = Directory(newPath);
    } else {
      await Permission.storage.request();
    }
  }
  if (!await directory!.exists()) {
    await directory.create(recursive: true);
  }
  if (await directory.exists()) {
    try {
      if (await _audioRecorder.hasPermission()) {
        final filePath =
            '${directory.path}/${DateTime.now().millisecondsSinceEpoch}.m4a';
        final isSupported = await _audioRecorder.isEncoderSupported(
          AudioEncoder.aacLc,
        );
        await _audioRecorder.start(path: filePath);
        _recordDuration = 0;
        _startTimer();
      }
    } catch (e) {
      print(e);
    }
  }
} catch (e) {
  print(e);
}
}

回答

需要更换包,record包的核心设计是将音频录制到本地文件,不提供实时音频流的获取能力,无法满足Azure认知语音服务实时转写的流式传输需求。推荐以下两种方案:

方案1:使用flutter_sound包获取实时音频流

flutter_sound支持实时获取麦克风的音频数据流,并且可以配置为Azure要求的WAV格式参数(16kHz采样率、16位深度、单声道)。核心代码示例:

import 'package:flutter_sound/flutter_sound.dart';

final FlutterSoundRecorder _recorder = FlutterSoundRecorder();
Stream<List<int>>? _audioStream;

Future<void> startStreaming() async {
  await _recorder.openRecorder();
  await _recorder.setSubscriptionDuration(const Duration(milliseconds: 100));
  
  // 配置音频参数匹配Azure要求
  _audioStream = _recorder.startRecorder(
    toStream: true,
    codec: Codec.pcm16WAV,
    sampleRate: 16000,
    numChannels: 1,
  );
  
  // 将_stream传入你的Azure API请求
  if (_audioStream != null) {
    sendStreamToAzure(_audioStream!);
  }
}

void sendStreamToAzure(Stream<List<int>> stream) {
  // 复用你已有的令牌获取和请求逻辑,替换getMicrophoneStream()为这个stream
}

方案2:使用Azure官方Flutter语音SDK

微软提供了官方的语音SDK包,直接封装了实时转写的全部逻辑,无需手动处理音频流和API调用,稳定性更高。核心示例:

import 'package:azure_cognitive_services_speech/azure_cognitive_services_speech.dart';

Future<void> startRealTimeRecognition() async {
  final speechConfig = SpeechConfig.fromSubscription(
    'YOUR_SUBSCRIPTION_KEY', 
    'YOUR_REGION'
  );
  speechConfig.speechRecognitionLanguage = 'en-US';
  
  final audioConfig = AudioConfig.fromDefaultMicrophoneInput();
  final recognizer = SpeechRecognizer(speechConfig, audioConfig);
  
  recognizer.recognized.listen((event) {
    if (event.result.reason == ResultReason.RecognizedSpeech) {
      print('转写结果: ${event.result.text}');
    }
  });
  
  await recognizer.startContinuousRecognitionAsync();
}

注意事项

  • 无论使用哪种方案,都需要确保音频参数与Azure要求一致:推荐16kHz采样率、16位PCM、单声道、WAV格式。
  • 实时转写需要保持网络连接稳定,建议添加断连重连逻辑。

内容的提问来源于stack exchange,提问作者thinkcode

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 17:45:02