Python文件处理问题:匹配特定内容仅读首行,如何保留所有匹配行?
问题:仅写入首行匹配内容,如何修改代码保留所有匹配行?
需求说明
读取指定目录下的目标文件,逐行搜索包含特定数字集4794 46的行,将所有匹配行拆分保存到新文件中。
文件示例
000002.4794 4600 9127 0320. 08/21 000001.4794 4600 0376 2836. 12/22
现有代码
for filename in path.glob(f"*{extension}"): with open(filename, encoding='utf-8') as infile: for line in infile: if '4794 46' in line: newpath = r'C:\RI\OUTBOX\479446' if not os.path.exists(newpath): os.makedirs(newpath) with open('479446//t'+now+'_visa_479446.dat', 'a', encoding='utf-8') as outfile: outfile.write(line)
问题现象
上述代码仅能写入第一行匹配4794 46的内容,后续匹配行被忽略。
解决方案
问题根源
原代码的写入逻辑被嵌套在if not os.path.exists(newpath)分支内,仅当目录不存在并创建时才会写入当前匹配行。后续匹配行触发时,目录已存在,不会进入该分支,因此无法写入。
修改要点
- 将目录创建逻辑移至循环外层,避免重复判断,提升效率
- 将文件写入逻辑移出目录判断分支,确保每一行匹配时都执行写入操作
- 使用
os.path.join拼接路径,避免硬编码路径分隔符,提升代码跨平台兼容性
修改后代码
import os from pathlib import Path # 提前定义输出路径和文件名 newpath = r'C:\RI\OUTBOX\479446' output_filename = os.path.join(newpath, f't{now}_visa_479446.dat') # 先创建目录(如果不存在) if not os.path.exists(newpath): os.makedirs(newpath) # 遍历目标文件 for filename in Path().glob(f"*{extension}"): with open(filename, encoding='utf-8') as infile: for line in infile: if '4794 46' in line: # 每次匹配行都打开文件追加写入 with open(output_filename, 'a', encoding='utf-8') as outfile: outfile.write(line)
优化建议
如果文件较大,频繁打开关闭文件会影响性能,可以将输出文件的打开逻辑移至外层,减少IO操作:
import os from pathlib import Path newpath = r'C:\RI\OUTBOX\479446' output_filename = os.path.join(newpath, f't{now}_visa_479446.dat') if not os.path.exists(newpath): os.makedirs(newpath) # 提前打开输出文件,全程保持打开状态 with open(output_filename, 'a', encoding='utf-8') as outfile: for filename in Path().glob(f"*{extension}"): with open(filename, encoding='utf-8') as infile: for line in infile: if '4794 46' in line: outfile.write(line)
内容的提问来源于stack exchange,提问作者Miko Efa
相关产品推荐
相关产品推荐

