Python分块读取文件触发UnicodeDecodeError编码报错该如何修复
分块读取UTF-8文件出现UnicodeDecodeError的修复方案
错误原因
UTF-8是变长编码,单个字符占用1~4字节不等,分块读取时切割边界刚好落在一个多字节字符的中间,导致当前块末尾的字节序列不是完整的UTF-8字符,直接解码就会触发unexpected end of data异常。一次性读取整个文件不存在切割问题,所以可以正常运行。
修复方案
方案1:使用增量解码器(推荐,无数据损失,代码简洁)
Python标准库codecs提供了增量解码器,专门用于分块解码场景,会自动处理跨块的不完整字符,无需手动处理异常。
对应修改后的代码如下:
import codecs fst = 0 # 初始化UTF-8增量解码器,若文件为其他编码替换即可,如gbk decoder = codecs.getincrementaldecoder('utf-8')() with open(in_ckfile, 'rb', 0) as file: with open(outfile_namepath, mode='wb') as outfile: while True: buf = file.read(204800) if not buf: # 读取结束,告知解码器输入完成,输出最后剩余内容 decoded_buf = decoder.decode(b'', final=True) if decoded_buf: xbytes = bytearray(map(ord, decoded_buf)) outfile.write(xbytes) break fst += 1 print('read no., len of buf ......: ', fst, len(buf)) decoded_buf = decoder.decode(buf) print('read no., len of decode buf: ', fst, len(decoded_buf)) xbytes = bytearray(map(ord, decoded_buf)) outfile.write(xbytes)
方案2:手动维护残余字节缓存
如果不想引入额外依赖,可以自己维护残余字节缓存,每次解码前拼接上一轮遗留的不完整字节,解码失败时将不完整的尾部字节留存到下一轮处理。
对应修改后的代码如下:
fst = 0 remain_bytes = b'' # 存储上一轮未解码的残余字节 with open(in_ckfile, 'rb', 0) as file: with open(outfile_namepath, mode='wb') as outfile: while True: buf = file.read(204800) if not buf: # 读取结束,处理最后残余字节 if remain_bytes: # 最后残余仍无法解码视为文件本身编码错误,可按需调整错误处理逻辑 decoded_buf = remain_bytes.decode(errors='replace') xbytes = bytearray(map(ord, decoded_buf)) outfile.write(xbytes) break fst += 1 print('read no., len of buf ......: ', fst, len(buf)) total_buf = remain_bytes + buf try: decoded_buf = total_buf.decode() remain_bytes = b'' except UnicodeDecodeError as e: # 解码到错误位置为止,错误位置后的字节留作下一轮残余 decoded_buf = total_buf[:e.start].decode() remain_bytes = total_buf[e.start:] print('read no., len of decode buf: ', fst, len(decoded_buf)) xbytes = bytearray(map(ord, decoded_buf)) outfile.write(xbytes)
注意事项
不推荐直接给decode方法添加errors='ignore'或errors='replace'参数跳过错误,该方法会直接丢弃不完整字符或替换为乱码,造成数据损失,仅在可接受数据失真的场景下使用。
内容的提问来源于stack exchange,提问作者user8597915
相关产品推荐
相关产品推荐

