You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dart读取大体积TXT文件报Unfinished UTF-8 octet sequence错误如何解决

报错根因

UTF-8为变长编码,单个字符占用1~4字节不等,固定块大小读取文件时,大概率会截断某个多字节字符的完整字节序列,导致读取到的字节数组末尾存在不完整的UTF-8八元组,直接调用utf8.decode解析就会触发该错误。

解决方案

提供两种可直接落地的实现方式:

方案1:手动处理剩余字节的逐块读取

适合需要精确控制读取偏移、块大小的场景,核心逻辑是每次解码前识别并暂存末尾不完整的UTF-8字节,拼接到下一次读取的字节数组头部再解析:

import 'dart:convert';
import 'dart:io';

void readFileInBlocks(String path, {int block = 64*64, int? skip}) async {
  final file = File(path);
  final raf = await file.open();
  if (skip != null) raf.setPositionSync(skip);
  // 存储上一块截断的残次字节
  List<int> remainingBytes = [];

  while (true) {
    final data = raf.readSync(block);
    if (data.isEmpty) break; // 读取完成
    // 拼接上一次的残次字节
    final fullBytes = [...remainingBytes, ...data];
    // 解码时自动处理不完整序列
    final decoder = Utf8Decoder(allowMalformed: false);
    String content = decoder.convert(fullBytes);
    // 计算末尾不完整字节的长度
    int remainingLength = _getUnfinishedUtf8Length(fullBytes);
    remainingBytes = fullBytes.sublist(fullBytes.length - remainingLength);
    // 按需更新content.value即可
    print(content);
  }
  raf.closeSync();
}

// 计算字节数组末尾不完整的UTF-8字节长度
int _getUnfinishedUtf8Length(List<int> bytes) {
  int length = 0;
  for (int i = bytes.length - 1; i >= 0; i--) {
    int b = bytes[i];
    if ((b & 0x80) == 0) break; // 单字节字符,无残次
    if ((b & 0xC0) == 0xC0) { // 多字节字符起始位
      int charLength = 0;
      if ((b & 0xE0) == 0xC0) charLength = 2;
      else if ((b & 0xF0) == 0xE0) charLength = 3;
      else if ((b & 0xF8) == 0xF0) charLength = 4;
      // 校验字符完整度
      if (bytes.length - i < charLength) {
        length = bytes.length - i;
      }
      break;
    }
    length++;
  }
  return length;
}

方案2:使用Dart内置文件流自动处理编码

如果不需要精确控制块读取逻辑,Dart的文件流内置了UTF-8编码的自动拼接处理,无需手动处理残次字节,代码更简洁,天然支持大文件读取:

import 'dart:convert';
import 'dart:io';

void readFileByStream(String path, int? skip) async {
  final file = File(path);
  // 传入start参数实现跳过指定字节逻辑,流式读取并自动解码UTF-8
  Stream<String> lines = file.openRead(skip ?? 0)
    .transform(utf8.decoder)
    .transform(LineSplitter());

  await for (String line in lines) {
    // 按行处理内容,按需更新content.value即可
    print(line);
  }
}

内容的提问来源于stack exchange,提问作者JasonZhao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 03:57:01