如何用Biopython修改FASTA文件:保留首个冒号前的序列ID
Using Biopython to Trim FASTA IDs
Absolutely! Biopython makes this kind of FASTA tweak totally straightforward. Here’s a solid solution that fits your exact input format (where sequences are on the same line as the headers):
from Bio import SeqIO # Replace these with your actual file paths input_file = "your_input.fasta" output_file = "trimmed_output.fasta" with open(output_file, "w") as out_handle: for record in SeqIO.parse(input_file, "fasta"): # Extract everything before the first colon as the new ID new_id = record.id.split(":")[0] # Handle cases where the sequence is on the same line as the header # If the record's built-in sequence is empty, use the description field (the part after the space) sequence = str(record.seq) if record.seq else record.description.strip() # Write the formatted line to the output file out_handle.write(f">{new_id} {sequence}\n")
How this works:
- We use
SeqIO.parse()to read through each record in your FASTA file. - The
split(':')[0]call takes the original ID (likeseq1:QXQXQWQ:XQWQ) and grabs just the part before the first colon, giving usseq1. - Since your input puts sequences on the same line as the headers, we check if the record's built-in sequence is empty (a quirk of how Biopython parses this format) and fall back to using the description field (which holds the sequence text in your case).
- Finally, we write each trimmed record in the exact format you want:
>new_id sequence.
Just plug in your actual file paths, run the script, and you’ll get the trimmed FASTA file you’re looking for.
内容的提问来源于stack exchange,提问作者Grendel
相关产品推荐
相关产品推荐

