Python中如何格式化FASTA格式蛋白质序列(合并多行)
Hey there, I get exactly what you're dealing with—collapsing those multi-line FASTA sequences into single lines is a super common first step before sequence alignment, and it's easy to hit snags with str.startswith if you don't handle line transitions just right. Regular expressions work great here, but I'll also share a more straightforward line-by-line approach in case you want something easier to debug later.
We can use regex to capture each FASTA header and all its corresponding sequence lines in one go, then clean up the sequence to remove newlines.
Here's a working example:
import re # Replace this with your actual FASTA content (or read from a file) fasta_data = """>A/Nayarit/InDRE2/2009 MDSGKGSSQKSRPKTRAPRRPSEPAVKRPVVEEPPAAEESV PPPEKRKPTTTASTQAETANPQVQPQPAAQPAPPVQPQAPQ PQAPQPAPQPAPQPAPQPQPQPAPQPAPQPQPQPAPQPAPQ PQPQPAPQPGAPQPQPQPPQPQPPQPQPQPQPQPQPQPQP >Another_Protein MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRIL PDGTVDGTRDRSDQHIQLQLSAESVGEVYIKSTETGQYLAM NLRGPYYIDGILKALNGGG """ # Regex pattern to match headers and their sequences pattern = r">(.*?)\n([^>]+)" matches = re.findall(pattern, fasta_data, re.DOTALL) # Build the formatted FASTA string formatted_fasta = "" for header, seq in matches: # Remove newlines and extra whitespace from the sequence cleaned_seq = seq.replace("\n", "").strip() formatted_fasta += f">{header}\n{cleaned_seq}\n" print(formatted_fasta)
How this works:
re.DOTALLmakes the.character match newlines, so we can capture multi-line sequences without breaking.([^>]+)grabs all content until the next>(the start of a new header), ensuring we get the full sequence for each entry.- We then strip out all newlines from the captured sequence to condense it into a single line.
If regex feels like a black box, this approach iterates through each line explicitly—it's great for catching edge cases like empty lines or unexpected formatting quirks:
def format_fasta(input_data): formatted_output = [] current_sequence = [] for line in input_data.splitlines(): line = line.strip() if not line: # Skip empty lines to avoid extra spaces continue if line.startswith(">"): # If we were building a sequence, add it to the output first if current_sequence: formatted_output.append("".join(current_sequence)) current_sequence = [] # Add the new header line to the output formatted_output.append(line) else: # Accumulate sequence lines to join later current_sequence.append(line) # Don't forget to add the last sequence in the file! if current_sequence: formatted_output.append("".join(current_sequence)) return "\n".join(formatted_output) # Example usage fasta_data = """>A/Nayarit/InDRE2/2009 MDSGKGSSQKSRPKTRAPRRPSEPAVKRPVVEEPPAAEESV PPPEKRKPTTTASTQAETANPQVQPQPAAQPAPPVQPQAPQ PQAPQPAPQPAPQPAPQPQPQPAPQPAPQPQPQPAPQPAPQ PQPQPAPQPGAPQPQPQPPQPQPPQPQPQPQPQPQPQPQP >Another_Protein MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRIL PDGTVDGTRDRSDQHIQLQLSAESVGEVYIKSTETGQYLAM NLRGPYYIDGILKALNGGG """ print(format_fasta(fasta_data))
How this works:
- We keep track of the current sequence being built as we loop through each line.
- When we hit a header line (
>), we finalize the previous sequence (if any) and start fresh for the new entry. - Sequence lines get added to a list, which we join into a single string once we hit the next header or reach the end of the file.
Reading from a File
If your FASTA is stored in a file instead of a string, just read it in first:
with open("your_proteins.fasta", "r") as file: fasta_data = file.read() # Then use either solution above formatted_result = format_fasta(fasta_data) # To save the formatted output back to a new file: with open("formatted_proteins.fasta", "w") as out_file: out_file.write(formatted_result)
内容的提问来源于stack exchange,提问作者Durian At

