Python匹配/Homer结尾模式后打印后续行直至下一个模式的实现方法
实现思路
- 先将
motif_list.txt预读取为字典,键为/Homer结尾的标识串,值为对应的motif类型,避免后续重复遍历查找,提升处理效率 - 遍历
annotation.txt每一行,维护当前生效的标识串和对应motif值:- 遇到/Homer结尾的行时,更新当前生效的标识和motif
- 遇到普通基因行时,用当前生效的标识和motif拼接输出
- 遇到分割线、空行直接跳过
完整可运行代码
import re # 读取motif_list为字典,实现O(1)速度查找对应关系 motif_dict = {} with open("motif_list.txt", "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue parts = line.split("\t") if len(parts) >= 2: key = parts[0].strip() motif_dict[key] = parts[1].strip() # 逐行处理annotation.txt current_homer_id = None current_motif_type = None with open("annotation.txt", "r", encoding="utf-8") as f: for line in f: raw_line = line.strip() # 跳过分割线和空行 if not raw_line or raw_line.startswith("---"): continue # 判断是否为/Homer结尾的标识行 if re.search(r"/Homer$", raw_line): if raw_line in motif_dict: current_homer_id = raw_line current_motif_type = motif_dict[raw_line] else: # 未匹配到的标识行,后续行不输出 current_homer_id = None current_motif_type = None continue # 是基因行且有生效的标识就按要求输出 if current_homer_id and current_motif_type: print(f"{raw_line}\t{current_homer_id}\t{current_motif_type}")
代码优化点
- 修复了原代码中
motif_into的拼写错误(应为motif_info) - 用字典替代嵌套循环查找motif对应关系,处理大文件时效率提升明显
- 统一对行做strip处理,避免换行符、首尾空格导致的匹配失败问题
- 增加了空行、分割线的跳过逻辑,避免无效内容输出
- 用
with语句管理文件句柄,执行完自动关闭文件,避免资源泄漏
内容的提问来源于stack exchange,提问作者Bitsy
相关产品推荐
相关产品推荐

