如何用Python去除FASTA基因组文件标题多余注释
处理FASTA文件标题,保留指定字段的Python实现
针对你的需求,这里提供一段适合新手理解的Python代码,能帮你批量处理FASTA文件,只保留标题中的locus_tag和protein字段,同时完整保留对应序列:
import re def process_fasta(input_file, output_file): # 定义匹配locus_tag和protein字段的正则表达式 locus_pattern = re.compile(r'\[locus_tag=([^\]]+)\]') protein_pattern = re.compile(r'\[protein=([^\]]+)\]') with open(input_file, 'r') as infile, open(output_file, 'w') as outfile: for line in infile: line = line.strip() if not line: continue # 跳过空行 # 判断是否是标题行(以>开头) if line.startswith('>'): # 提取标题内容(去掉开头的>) header = line[1:] # 查找locus_tag和protein字段 locus_match = locus_pattern.search(header) protein_match = protein_pattern.search(header) # 构建新标题,确保字段存在才添加 new_header_parts = [] if locus_match: new_header_parts.append(f'[locus_tag={locus_match.group(1)}]') if protein_match: new_header_parts.append(f'[protein={protein_match.group(1)}]') # 拼接新标题,开头加> new_header = '>' + ' '.join(new_header_parts) outfile.write(new_header + '\n') else: # 序列行直接写入 outfile.write(line + '\n') # 使用示例:替换成你的输入输出文件路径 process_fasta('input.fasta', 'output.fasta')
代码说明:
- 正则表达式匹配:用
re.compile预编译匹配规则,[^\]]+表示匹配除了]之外的所有字符,精准提取字段内容。 - 文件读写:用
with语句自动管理文件打开/关闭,避免资源泄漏。 - 标题处理逻辑:
- 识别以
>开头的标题行,去掉开头的>后提取内容。 - 分别查找
locus_tag和protein字段,存在则加入新标题。 - 拼接新标题并写入输出文件。
- 识别以
- 序列处理:非标题行直接原样写入,保证序列完整。
效果示例:
输入标题:
lcl|CP000438.1_cds_ABJ14958.1_2 [gene=dnaN] [locus_tag=PA14_00020] [protein=DNA polymerase III, beta chain] [protein_id=ABJ14958.1] [location=2056..3159] [gbkey=CDS]
输出标题:
[locus_tag=PA14_00020] [protein=DNA polymerase III, beta chain]
序列部分完全保留原内容。
内容的提问来源于stack exchange,提问作者Gracious Donkor
相关产品推荐
相关产品推荐

