You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取含BOM与CRLF的文本文件并去除BOM字符?

如何读取带BOM的UTF-8文本文件并自动去除BOM?

问题背景

我有一批带字节顺序标记(BOM)、使用CRLF(\r\n)作为换行符的UTF-8文本文件,八进制转储片段如下:

$ od -bc P21_T_3-28-2022.txt
0000000   357 273 277 163 164 141 147 145 040 061 015 012 120 154 141 171
         357 273 277   s   t   a   g   e       1  \r  \n   P   l   a   y
0000020   151 156 147 040 164 150 145 163 145 040 164 167 157 040 147 141
           i   n   g       t   h   e   s   e       t   w   o       g   a
0000040   155 145 163 054 040 162 145 155 151 156 144 145 144 040 155 145
           m   e   s   ,       r   e   m   i   n   d   e   d       m   e
0000060   040 157 146 040 164 151 155 145 163 040 164 150 141 164 040 111
               o   f       t   i   m   e   s       t   h   a   t       I
<snip>

用以下代码读取时,第一行输出会带有BOM字符\ufeff:

lines = open(file, "r", encoding='utf-8').read().splitlines()
print(lines[0])  # 输出: '\ufeffstage 1'

解决方案

方法1:使用utf-8-sig编码(推荐)

Python的utf-8-sig编码会自动识别并移除UTF-8文件开头的BOM字符,直接替换编码参数即可:

lines = open(file, "r", encoding='utf-8-sig').read().splitlines()
print(lines[0])  # 输出: 'stage 1'

这是官方推荐的处理带BOM UTF-8文件的方式,无需手动处理,简洁可靠。

方法2:手动去除BOM(适用于必须使用utf-8编码的场景)

如果因为某些原因不能切换编码,可以读取内容后手动检查并移除开头的BOM:

with open(file, "r", encoding='utf-8') as f:
    content = f.read()

# 检查并移除开头的BOM字符
if content.startswith('\ufeff'):
    content = content[1:]

lines = content.splitlines()
print(lines[0])  # 输出: 'stage 1'

内容的提问来源于stack exchange,提问作者codeforester

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 08:18:20