You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python分块读取文件触发UnicodeDecodeError编码报错该如何修复

分块读取UTF-8文件出现UnicodeDecodeError的修复方案

错误原因

UTF-8是变长编码,单个字符占用1~4字节不等,分块读取时切割边界刚好落在一个多字节字符的中间,导致当前块末尾的字节序列不是完整的UTF-8字符,直接解码就会触发unexpected end of data异常。一次性读取整个文件不存在切割问题,所以可以正常运行。

修复方案

方案1:使用增量解码器(推荐,无数据损失,代码简洁)

Python标准库codecs提供了增量解码器,专门用于分块解码场景,会自动处理跨块的不完整字符,无需手动处理异常。
对应修改后的代码如下:

import codecs

fst = 0
# 初始化UTF-8增量解码器,若文件为其他编码替换即可,如gbk
decoder = codecs.getincrementaldecoder('utf-8')()
with open(in_ckfile, 'rb', 0) as file:
    with open(outfile_namepath, mode='wb') as outfile:
        while True:
            buf = file.read(204800)
            if not buf:
                # 读取结束,告知解码器输入完成,输出最后剩余内容
                decoded_buf = decoder.decode(b'', final=True)
                if decoded_buf:
                    xbytes = bytearray(map(ord, decoded_buf))
                    outfile.write(xbytes)
                break
                    
            fst += 1
            print('read no., len of buf ......: ', fst, len(buf))

            decoded_buf = decoder.decode(buf)
            print('read no., len of decode buf: ', fst, len(decoded_buf))
            
            xbytes = bytearray(map(ord, decoded_buf))
            outfile.write(xbytes)

方案2:手动维护残余字节缓存

如果不想引入额外依赖,可以自己维护残余字节缓存,每次解码前拼接上一轮遗留的不完整字节,解码失败时将不完整的尾部字节留存到下一轮处理。
对应修改后的代码如下:

fst = 0
remain_bytes = b''  # 存储上一轮未解码的残余字节
with open(in_ckfile, 'rb', 0) as file:
    with open(outfile_namepath, mode='wb') as outfile:
        while True:
            buf = file.read(204800)
            if not buf:
                # 读取结束,处理最后残余字节
                if remain_bytes:
                    # 最后残余仍无法解码视为文件本身编码错误,可按需调整错误处理逻辑
                    decoded_buf = remain_bytes.decode(errors='replace')
                    xbytes = bytearray(map(ord, decoded_buf))
                    outfile.write(xbytes)
                break
                    
            fst += 1
            print('read no., len of buf ......: ', fst, len(buf))

            total_buf = remain_bytes + buf
            try:
                decoded_buf = total_buf.decode()
                remain_bytes = b''
            except UnicodeDecodeError as e:
                # 解码到错误位置为止,错误位置后的字节留作下一轮残余
                decoded_buf = total_buf[:e.start].decode()
                remain_bytes = total_buf[e.start:]

            print('read no., len of decode buf: ', fst, len(decoded_buf))
            
            xbytes = bytearray(map(ord, decoded_buf))
            outfile.write(xbytes)

注意事项

不推荐直接给decode方法添加errors='ignore'或errors='replace'参数跳过错误,该方法会直接丢弃不完整字符或替换为乱码,造成数据损失,仅在可接受数据失真的场景下使用。

内容的提问来源于stack exchange,提问作者user8597915

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 13:36:03