Python实现:移除文本行开头匹配的前两个字符'11'
移除文本每行开头的'11'字符解决方案
问题场景
我有一个较大的文本文件file.txt,内容格式如下,希望移除每行开头的两个字符11:
11112345,67890,12345 115432,a123q,hs1230 11s1a123,qw321,98765321 342342,121sa,12123243 11023456,sa123,d32acas2
我的代码(存在问题)
import re with open('in.txt') as oldfile, open('out.txt', 'w') as newfile: for line in oldfile: removed = re.sub(r'11', '', line[:2]): # 语法错误+逻辑缺陷 newfile.write(removed)
期望结果
112345,67890,12345 115432,a123q,hs1230 s1a123,qw321,98765321 342342,121sa,12123243 023456,sa123,d32acas2
修正方案
- 方法1:字符串切片(高效,适合固定开头匹配)
无需正则,直接判断每行开头是否为11,是则从第3个字符截取,否则保留原行:
with open('file.txt', 'r') as oldfile, open('out.txt', 'w') as newfile: for line in oldfile: # 先处理换行符避免干扰,处理后还原 stripped_line = line.rstrip('\n') if stripped_line.startswith('11'): new_line = stripped_line[2:] + '\n' else: new_line = line newfile.write(new_line)
- 方法2:正则表达式(适合灵活匹配场景)
用^11锚定行开头,仅替换开头的11,预编译正则提升大文件处理效率:
import re with open('file.txt', 'r') as oldfile, open('out.txt', 'w') as newfile: pattern = re.compile(r'^11') for line in oldfile: new_line = pattern.sub('', line) newfile.write(new_line)
原代码问题说明
- 语法错误:赋值语句末尾多了冒号,导致代码无法运行
- 逻辑错误:仅处理了每行前2个字符,完全丢弃了后面的内容
- 匹配缺陷:正则未锚定开头,且未考虑开头不是
11的行,处理逻辑不完整
内容的提问来源于stack exchange,提问作者kng
相关产品推荐
相关产品推荐

