You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则分割NCBI-BLAST结果并存入列表分析

用Python正则表达式分割NCBI BLAST输出并存储到列表

我经常帮人处理BLAST输出的解析问题,你的需求很典型——把每条显著对齐的结果拆出来存成列表方便后续分析。先结合你给的示例文本,给你两种解决方案:一种是用正则表达式手动解析,另一种是用专业生物信息学库更省心的方式。


方案一:正则表达式手动解析(针对文本格式输出)

先看你的示例输出结构,关键是要定位到Sequences producing significant alignments:之后的每条结果行。我写了一段代码,直接可以复用:

import re

# 替换成你的实际BLAST输出文本(或者从文件读取)
blast_output = """Query= lcl|TRINITY_DN2888_c0_g2_i1 Length=1394
Score E
Sequences producing significant alignments:
(Bits) Value
sp|Q9S775|PKL_ARATH CHD3-type chromatin-remodeling factor PICKLE... 1640 0.0
sp|Q9S775|PKL_ARATH CHD3-type chromatin-remodeling factor P..."""

# 构建正则模式,匹配每条对齐条目
# 适配以sp|开头的Swiss-Prot条目,其他数据库前缀可以自行扩展
alignment_pattern = re.compile(
    r'^(?P<accession>sp\|[^\s]+)\s+(?P<description>.+?)\s+(?P<score_bits>\d+)\s+(?P<e_value>[\d\.]+)$',
    re.MULTILINE
)

# 解析所有条目并存入列表
alignment_list = []
for match in alignment_pattern.finditer(blast_output):
    # 把每条结果转成字典,方便后续提取字段
    alignment_entry = match.groupdict()
    alignment_list.append(alignment_entry)

# 验证结果
for idx, entry in enumerate(alignment_list, 1):
    print(f"条目 {idx}:")
    print(f"  登录号: {entry['accession']}")
    print(f"  描述: {entry['description']}")
    print(f"  得分(Bits): {entry['score_bits']}")
    print(f"  E值: {entry['e_value']}\n")

代码说明:

  • 用re.MULTILINE让^匹配每行开头,确保只抓每条结果的行。
  • 正则里用了命名捕获组((?P<name>)),直接转成字典结构,后续分析时提取字段更方便。
  • 如果你的输出里还有其他数据库的条目(比如gb|、ref|),可以把正则开头改成^(?P<accession>(sp\|gb\|ref\|)[^\s]+)来兼容。

方案二:用BioPython库解析(更可靠,推荐)

如果你的BLAST输出是XML或者TAB格式,强烈推荐用BioPython的BLAST解析模块,不用自己写正则,避免格式变动导致匹配失败的问题:

from Bio.Blast import NCBIXML

# 读取XML格式的BLAST输出文件
with open("your_blast_output.xml", "r") as blast_file:
    blast_records = NCBIXML.parse(blast_file)
    
    alignment_list = []
    for record in blast_records:
        # 遍历每条查询序列的对齐结果
        for alignment in record.alignments:
            # 遍历每条对齐的高得分片段(HSP)
            for hsp in alignment.hsps:
                alignment_entry = {
                    "query_id": record.query_id,
                    "accession": alignment.accession,
                    "description": alignment.title,
                    "score_bits": hsp.bits,
                    "e_value": hsp.expect,
                    "alignment_length": hsp.align_length
                }
                alignment_list.append(alignment_entry)

# 现在alignment_list里就存好了所有结构化的对齐结果

为什么推荐这个方案?

BLAST的文本格式偶尔会有换行或者格式变动,正则很容易失效,但BioPython是专门为生物信息学数据设计的,能完美兼容官方输出格式,后续扩展分析(比如提取序列比对片段、统计得分)也更方便。


内容的提问来源于stack exchange,提问作者Jithin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:49:16