You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用re.findall迭代SeqIO创建的FASTA序列列表?

问题解决:处理FASTA文件的多模式匹配

我在这个问题上已经花了10多个小时。输入是FASTA格式文件,需要输出包含基因ID与三种匹配模式的文本文件。原本想写自定义函数避免重复代码但没成功,只能重复三次代码实现功能。现在想用records = list(SeqIO.parse('mytextfile.fasta', 'fasta'))替代重复代码,用到Bio和re模块,但尝试的代码报错“findall()缺少必填参数'string'”,不知道正确传参方式。


现有可行代码

from Bio import SeqIO
import re

outfile = 'sekvenser.txt'

for seq_record in SeqIO.parse('prot_sequences.fasta', 'fasta'):
    match = re.findall(r'W.P', str(seq_record.seq), re.I)
    if match:       
        with open(outfile, 'a') as f:
            record_string = str(seq_record.id)
            newmatch = str(match)
            result = record_string+'\t'+newmatch
            print(result)
            f.write(result + '\n')

尝试的错误代码

records = list(SeqIO.parse('prot_sequences.fasta', 'fasta'))
new_list = []
i = r'W.P'

for i in records:
    match = re.findall(i)
    if match:       
        new_list.append(match)

print(new_list)

运行后报错:findall() is missing 1 required positional argument: 'string'。我以为i是字符串,但显然操作有误。如果尝试插入之前用的seq_record,会提示seq_record未定义,不知道该在i之后传什么参数。


问题分析与修复

错误根源有两点:

  1. 变量名冲突:先将正则模式赋值给i,后续循环又用for i in records,把i覆盖成了SeqIO的序列对象,不再是正则表达式字符串。
  2. 参数缺失:re.findall()的正确用法是re.findall(正则模式, 要匹配的字符串, 可选标志),你只传了第一个参数,没传入要匹配的序列字符串。

优化后的多模式匹配代码

from Bio import SeqIO
import re

# 定义三种需要匹配的正则模式
patterns = [r'W.P', r'你的第二种模式', r'你的第三种模式']
outfile = 'sekvenser.txt'
# 一次性读取所有FASTA序列记录
records = list(SeqIO.parse('prot_sequences.fasta', 'fasta'))

# 先清空输出文件,避免多次运行内容重复
with open(outfile, 'w') as f:
    pass

# 遍历所有匹配模式
for pattern in patterns:
    # 遍历每个序列记录
    for seq_record in records:
        seq_str = str(seq_record.seq)
        # 执行正则匹配,忽略大小写
        matches = re.findall(pattern, seq_str, re.I)
        if matches:
            result = f"{seq_record.id}\t{matches}"
            print(result)
            # 追加写入结果到文件
            with open(outfile, 'a') as f:
                f.write(result + '\n')

代码说明

  • 将三种正则模式存入patterns列表,通过循环遍历实现多模式匹配,彻底避免重复代码。
  • 用records = list(SeqIO.parse(...))一次性读取所有序列,后续可重复调用,提升效率。
  • 严格按照re.findall()的参数要求传参:正则模式、序列字符串、忽略大小写标志。
  • 先以w模式打开输出文件清空内容,再用a模式追加结果,避免多次运行时内容冗余。

内容的提问来源于stack exchange,提问作者Sara Stensson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 23:53:19