如何读取含BOM与CRLF的文本文件并去除BOM字符?
如何读取带BOM的UTF-8文本文件并自动去除BOM?
问题背景
我有一批带字节顺序标记(BOM)、使用CRLF(\r\n)作为换行符的UTF-8文本文件,八进制转储片段如下:
$ od -bc P21_T_3-28-2022.txt 0000000 357 273 277 163 164 141 147 145 040 061 015 012 120 154 141 171 357 273 277 s t a g e 1 \r \n P l a y 0000020 151 156 147 040 164 150 145 163 145 040 164 167 157 040 147 141 i n g t h e s e t w o g a 0000040 155 145 163 054 040 162 145 155 151 156 144 145 144 040 155 145 m e s , r e m i n d e d m e 0000060 040 157 146 040 164 151 155 145 163 040 164 150 141 164 040 111 o f t i m e s t h a t I <snip>
用以下代码读取时,第一行输出会带有BOM字符\ufeff:
lines = open(file, "r", encoding='utf-8').read().splitlines() print(lines[0]) # 输出: '\ufeffstage 1'
解决方案
方法1:使用utf-8-sig编码(推荐)
Python的utf-8-sig编码会自动识别并移除UTF-8文件开头的BOM字符,直接替换编码参数即可:
lines = open(file, "r", encoding='utf-8-sig').read().splitlines() print(lines[0]) # 输出: 'stage 1'
这是官方推荐的处理带BOM UTF-8文件的方式,无需手动处理,简洁可靠。
方法2:手动去除BOM(适用于必须使用utf-8编码的场景)
如果因为某些原因不能切换编码,可以读取内容后手动检查并移除开头的BOM:
with open(file, "r", encoding='utf-8') as f: content = f.read() # 检查并移除开头的BOM字符 if content.startswith('\ufeff'): content = content[1:] lines = content.splitlines() print(lines[0]) # 输出: 'stage 1'
内容的提问来源于stack exchange,提问作者codeforester
相关产品推荐
相关产品推荐

