如何用Python删除TXT文件中'PROCESSING'至'---'间的内容?
解决方案
你之前的代码逻辑是筛选行首为字母且后续两位是数字的行,和你要删除特定区块的需求完全不匹配,这是没达到效果的核心原因。针对你的需求,分两种场景给出实现:
场景1:删除单行中同时以PROCESSING开头、以---结尾的内容
如果目标内容是单行(比如某行完整是PROCESSINGxxxx---),可以直接用startswith()和endswith()判断:
with open('/Users/sebastian/Desktop/terms codes script clean/firmscodes.txt', 'r') as infile, open('test.txt', 'w') as outfile: for line in infile: # 去除换行符影响后再判断结尾 if not (line.startswith('PROCESSING') and line.rstrip('\n').endswith('---')): outfile.write(line)
场景2:删除多行区块(从PROCESSING开头的行到---结尾的行的所有内容)
如果是多行的区块(比如从某行PROCESSING...开始,中间若干行,直到某行...---结束的整个块),需要跟踪是否处于待删除的区块内,这种方式适合处理大型文件(无需一次性加载所有行到内存):
with open('/Users/sebastian/Desktop/terms codes script clean/firmscodes.txt', 'r') as infile, open('test.txt', 'w') as outfile: in_delete_block = False for line in infile: if not in_delete_block: # 未进入删除块,检查是否是区块起始行 if line.startswith('PROCESSING'): in_delete_block = True else: outfile.write(line) else: # 处于删除块内,检查是否是区块结束行 if line.rstrip('\n').endswith('---'): in_delete_block = False # 块内所有行直接跳过,不写入
关键说明
- 用
rstrip('\n')处理行尾换行符,避免因为换行符存在导致endswith('---')匹配失败。 - 逐行处理的方式内存占用低,适合GB级别的大型文本文件。
内容的提问来源于stack exchange,提问作者Sebastian Peralta
相关产品推荐
相关产品推荐

