You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Biopython修改FASTA文件:保留首个冒号前的序列ID

Using Biopython to Trim FASTA IDs

Absolutely! Biopython makes this kind of FASTA tweak totally straightforward. Here’s a solid solution that fits your exact input format (where sequences are on the same line as the headers):

from Bio import SeqIO

# Replace these with your actual file paths
input_file = "your_input.fasta"
output_file = "trimmed_output.fasta"

with open(output_file, "w") as out_handle:
    for record in SeqIO.parse(input_file, "fasta"):
        # Extract everything before the first colon as the new ID
        new_id = record.id.split(":")[0]
        
        # Handle cases where the sequence is on the same line as the header
        # If the record's built-in sequence is empty, use the description field (the part after the space)
        sequence = str(record.seq) if record.seq else record.description.strip()
        
        # Write the formatted line to the output file
        out_handle.write(f">{new_id} {sequence}\n")

How this works:

  • We use SeqIO.parse() to read through each record in your FASTA file.
  • The split(':')[0] call takes the original ID (like seq1:QXQXQWQ:XQWQ) and grabs just the part before the first colon, giving us seq1.
  • Since your input puts sequences on the same line as the headers, we check if the record's built-in sequence is empty (a quirk of how Biopython parses this format) and fall back to using the description field (which holds the sequence text in your case).
  • Finally, we write each trimmed record in the exact format you want: >new_id sequence.

Just plug in your actual file paths, run the script, and you’ll get the trimmed FASTA file you’re looking for.

内容的提问来源于stack exchange,提问作者Grendel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:33:16