You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python使用SeqIO处理FASTA文件重复序列ID输出格式调整问题

代码修改方案

原有代码每次遍历到重复序列对应的记录就单独调用print(),print()默认会在输出结尾添加换行符,因此每个ID单独占一行,调整打印逻辑即可实现预期效果。

修改后的完整代码

from Bio import SeqIO

def seqmatchsequence(raw_input):
    records = list(SeqIO.parse(raw_input, "fasta"))
    d = dict()
    for record in records:
        if record.seq in d:
            d[record.seq].append(record)
        else:
            d[record.seq] = [record]
    for seq, record_set in d.items():
        record_count = len(record_set)
        if record_count != 1:
            print(f': ({record_count})')
            # 提取所有ID拼接后一次性打印
            print(', '.join([record.id for record in record_set]))

核心修改点

  • 提前统计重复序列对应的记录数量,避免重复计算
  • 用列表推导式批量提取所有重复记录的ID,使用, 作为分隔符拼接为单个字符串
  • 仅调用一次print()输出所有ID,避免自动换行分割内容

内容的提问来源于stack exchange,提问作者Gerard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 16:06:04