UTF-8无BOM格式TXT文件处理异常求助:Python代码运行问题
UTF-8无BOM大文件处理乱码解决方案
问题根源
你当前的代码未明确指定文件编码,Python的io.open在不同操作系统下默认编码存在差异(例如Windows默认采用GBK编码),导致UTF-8无BOM格式的源文件被错误解码,写入时又用错误编码保存,最终出现乱码。6GB大文件的处理逻辑本身没问题,但编码匹配是核心问题。
修正后的代码
import time start = time.time() print(f"Empece en: {start}") # 使用with语句同时管理输入输出文件,自动处理资源释放 with open("Edificado/edificco.txt", "r", encoding="utf-8", errors="ignore") as input_file, \ open("Edificado/edificadoResultadoSinPorBlancos.txt", "w", encoding="utf-8") as output_file: for line in input_file: if '|' in line: line = line.replace("|", "-") output_file.write(line) end = time.time() print(f"Termine en: {end - start}")
核心要点
- 强制指定编码:读、写文件时都明确设置
encoding='utf-8',确保源文件的UTF-8无BOM编码被正确解析,输出文件也保持一致编码 - 优化文件管理:用
with语句同时管控两个文件,无需手动调用close(),避免资源泄漏 - 大文件适配:逐行处理的方式不会加载整个6GB文件到内存,保持内存占用稳定
内容的提问来源于stack exchange,提问作者imnachox2
相关产品推荐
相关产品推荐

