You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flutter使用record包处理PCM16音频流:是否需添加流头部?

问题解答:是否需要为PCM16音频流添加头部?

需要。你当前录制的是裸PCM16格式的音频数据,这类数据仅包含原始音频采样值,没有任何描述音频属性的元信息(比如采样率、通道数、位深)。语音转文本服务必须依赖这些元信息才能正确解析音频内容,直接发送裸PCM数据会导致服务无法识别或解析错误。

为什么需要添加头部?

WAV格式本质是「PCM数据 + 文件头」的组合,文件头包含以下关键信息:

  • 采样率(语音转文本通常用16000Hz)
  • 通道数(语音识别一般用单声道)
  • 位深(这里为16bit)
  • 音频数据总长度

缺少这些信息,后端服务无法判断如何解码原始采样数据,自然无法完成语音转文本。

具体实现建议

1. 生成标准WAV文件头

在录制完成后,先为audioDataBuffer添加WAV文件头,再发送给后端。以下是适配PCM16单声道的WAV头生成示例:

List<int> generateWavHeader(int sampleRate, int channels, int bitDepth, int dataLength) {
  final int fileSize = 36 + dataLength; // WAV头固定36字节 + PCM数据长度
  final List<int> header = List<int>.filled(44, 0);

  // RIFF标识
  header[0] = 'R'.codeUnitAt(0);
  header[1] = 'I'.codeUnitAt(0);
  header[2] = 'F'.codeUnitAt(0);
  header[3] = 'F'.codeUnitAt(0);
  // 文件总大小
  header[4] = fileSize & 0xff;
  header[5] = (fileSize >> 8) & 0xff;
  header[6] = (fileSize >> 16) & 0xff;
  header[7] = (fileSize >> 24) & 0xff;
  // WAVE标识
  header[8] = 'W'.codeUnitAt(0);
  header[9] = 'A'.codeUnitAt(0);
  header[10] = 'V'.codeUnitAt(0);
  header[11] = 'E'.codeUnitAt(0);
  // fmt子块标识
  header[12] = 'f'.codeUnitAt(0);
  header[13] = 'm'.codeUnitAt(0);
  header[14] = 't'.codeUnitAt(0);
  header[15] = ' '.codeUnitAt(0);
  // fmt子块长度(PCM格式固定为16)
  header[16] = 16;
  header[17] = 0;
  header[18] = 0;
  header[19] = 0;
  // 音频格式(PCM为1)
  header[20] = 1;
  header[21] = 0;
  // 通道数
  header[22] = channels;
  header[23] = 0;
  // 采样率
  header[24] = sampleRate & 0xff;
  header[25] = (sampleRate >> 8) & 0xff;
  header[26] = (sampleRate >> 16) & 0xff;
  header[27] = (sampleRate >> 24) & 0xff;
  // 字节率 = 采样率 * 通道数 * 位深/8
  final int byteRate = sampleRate * channels * (bitDepth ~/ 8);
  header[28] = byteRate & 0xff;
  header[29] = (byteRate >> 8) & 0xff;
  header[30] = (byteRate >> 16) & 0xff;
  header[31] = (byteRate >> 24) & 0xff;
  // 块对齐 = 通道数 * 位深/8
  final int blockAlign = channels * (bitDepth ~/ 8);
  header[32] = blockAlign & 0xff;
  header[33] = (blockAlign >> 8) & 0xff;
  // 位深
  header[34] = bitDepth;
  header[35] = 0;
  // data子块标识
  header[36] = 'd'.codeUnitAt(0);
  header[37] = 'a'.codeUnitAt(0);
  header[38] = 't'.codeUnitAt(0);
  header[39] = 'a'.codeUnitAt(0);
  // PCM数据长度
  header[40] = dataLength & 0xff;
  header[41] = (dataLength >> 8) & 0xff;
  header[42] = (dataLength >> 16) & 0xff;
  header[43] = (dataLength >> 24) & 0xff;

  return header;
}

2. 修改发送逻辑

在sendFileToBackend中,合并WAV头与PCM数据后再创建MultipartFile:

void sendFileToBackend() async {
  if(audioDataBuffer.isNotEmpty){
    // 匹配录制配置:16000Hz采样率、单声道、16bit位深
    final wavHeader = generateWavHeader(16000, 1, 16, audioDataBuffer.length);
    final fullWavData = [...wavHeader, ...audioDataBuffer];

    Uri uri = Uri.parse('http://localhost:8000/speech-to-text/');
    var request = http.MultipartRequest('POST', uri);
    var file = http.MultipartFile.fromBytes(
        'file', fullWavData, filename: 'speech.wav',
        contentType: MediaType('audio', 'wav')); // 改为标准audio/wav类型
    request.files.add(file);
    // 发送请求并处理响应
    final response = await request.send();
    // 补充响应处理逻辑...
  }
}

注意事项

  • 确保generateWavHeader的参数(采样率、通道数)与RecordConfig中的配置一致,比如在录制时显式设置:RecordConfig(encoder: AudioEncoder.pcm16bits, sampleRate: 16000, numChannels: 1, ...)
  • 如果需要实时流式发送(边录边发),WAV头的总数据长度无法提前确定,这种情况可发送裸PCM,但需在请求头中告知后端采样率、通道数等参数,具体需遵循后端API要求。

内容的提问来源于stack exchange,提问作者Steffen Kämmerer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 16:43:29