You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取FTP导出的固定宽度文件Windows正常Linux报错如何解决?

问题根因

  • encoding=None 参数的实际效果依赖操作系统默认编码:Windows默认编码为cp1252/GBK类编码,恰好兼容源文件中的0x99字节,所以运行正常;Linux系统默认编码为UTF-8,无法解析该字节才抛出UnicodeDecodeError
  • 固定宽度文件的字段长度是按字节计数,不需要解码为字符串处理,使用文本模式读写反而会引入编码兼容问题

修复方案

方案1:二进制模式读写(最稳妥,无需考虑编码)

全程按字节处理字段,完全规避编码报错:

with open("recode.dat", "rb") as open_src:
    with open("target_file.dat", "wb+") as open_tgt:
        for src_rec in open_src:
            new_rec = b''
            for f_length in data_type_length:
                f_length = int(f_length)
                field = b'"' + src_rec[:f_length].strip() + b'"|'
                new_rec += field
                src_rec = src_rec[f_length:]
            open_tgt.write(new_rec[:-1] + b'\n')

方案2:指定通用兼容编码

如果需要转为字符串做额外处理,可指定latin-1编码,该编码兼容所有单字节字符,不会抛出解码错误:

with open("recode.dat", "r", encoding="latin-1", errors="replace") as open_src:
    with open("target_file.dat", "w+", encoding="latin-1") as open_tgt:
        # 原有字符串处理逻辑保持不变即可
        for src_rec in open_src:
            new_rec = ''
            for f_length in data_type_length:
                f_length = int(f_length)
                field = '"' + src_rec[:f_length].strip() + '"|'
                new_rec += (field)
                src_rec = src_rec[f_length:]
            open_tgt.write(new_rec[:-1] + '\n')

内容的提问来源于stack exchange,提问作者Yohan Neranga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 19:54:04