You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

UTF-8无BOM格式TXT文件处理异常求助:Python代码运行问题

UTF-8无BOM大文件处理乱码解决方案

问题根源

你当前的代码未明确指定文件编码,Python的io.open在不同操作系统下默认编码存在差异(例如Windows默认采用GBK编码),导致UTF-8无BOM格式的源文件被错误解码,写入时又用错误编码保存,最终出现乱码。6GB大文件的处理逻辑本身没问题,但编码匹配是核心问题。

修正后的代码

import time

start = time.time()
print(f"Empece en: {start}")

# 使用with语句同时管理输入输出文件,自动处理资源释放
with open("Edificado/edificco.txt", "r", encoding="utf-8", errors="ignore") as input_file, \
     open("Edificado/edificadoResultadoSinPorBlancos.txt", "w", encoding="utf-8") as output_file:
    for line in input_file:
        if '|' in line:
            line = line.replace("|", "-")
        output_file.write(line)

end = time.time()
print(f"Termine en: {end - start}")

核心要点

  • 强制指定编码:读、写文件时都明确设置encoding='utf-8',确保源文件的UTF-8无BOM编码被正确解析,输出文件也保持一致编码
  • 优化文件管理:用with语句同时管控两个文件,无需手动调用close(),避免资源泄漏
  • 大文件适配:逐行处理的方式不会加载整个6GB文件到内存,保持内存占用稳定

内容的提问来源于stack exchange,提问作者imnachox2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 02:40:29