Python读取FTP导出的固定宽度文件Windows正常Linux报错如何解决?
问题根因
encoding=None参数的实际效果依赖操作系统默认编码:Windows默认编码为cp1252/GBK类编码,恰好兼容源文件中的0x99字节,所以运行正常;Linux系统默认编码为UTF-8,无法解析该字节才抛出UnicodeDecodeError- 固定宽度文件的字段长度是按字节计数,不需要解码为字符串处理,使用文本模式读写反而会引入编码兼容问题
修复方案
方案1:二进制模式读写(最稳妥,无需考虑编码)
全程按字节处理字段,完全规避编码报错:
with open("recode.dat", "rb") as open_src: with open("target_file.dat", "wb+") as open_tgt: for src_rec in open_src: new_rec = b'' for f_length in data_type_length: f_length = int(f_length) field = b'"' + src_rec[:f_length].strip() + b'"|' new_rec += field src_rec = src_rec[f_length:] open_tgt.write(new_rec[:-1] + b'\n')
方案2:指定通用兼容编码
如果需要转为字符串做额外处理,可指定latin-1编码,该编码兼容所有单字节字符,不会抛出解码错误:
with open("recode.dat", "r", encoding="latin-1", errors="replace") as open_src: with open("target_file.dat", "w+", encoding="latin-1") as open_tgt: # 原有字符串处理逻辑保持不变即可 for src_rec in open_src: new_rec = '' for f_length in data_type_length: f_length = int(f_length) field = '"' + src_rec[:f_length].strip() + '"|' new_rec += (field) src_rec = src_rec[f_length:] open_tgt.write(new_rec[:-1] + '\n')
内容的提问来源于stack exchange,提问作者Yohan Neranga
相关产品推荐
相关产品推荐

