You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何格式化FASTA格式蛋白质序列(合并多行)

Hey there, I get exactly what you're dealing with—collapsing those multi-line FASTA sequences into single lines is a super common first step before sequence alignment, and it's easy to hit snags with str.startswith if you don't handle line transitions just right. Regular expressions work great here, but I'll also share a more straightforward line-by-line approach in case you want something easier to debug later.

Solution 1: Regular Expression (Quick & Concise)

We can use regex to capture each FASTA header and all its corresponding sequence lines in one go, then clean up the sequence to remove newlines.

Here's a working example:

import re

# Replace this with your actual FASTA content (or read from a file)
fasta_data = """>A/Nayarit/InDRE2/2009
MDSGKGSSQKSRPKTRAPRRPSEPAVKRPVVEEPPAAEESV
PPPEKRKPTTTASTQAETANPQVQPQPAAQPAPPVQPQAPQ
PQAPQPAPQPAPQPAPQPQPQPAPQPAPQPQPQPAPQPAPQ
PQPQPAPQPGAPQPQPQPPQPQPPQPQPQPQPQPQPQPQP
>Another_Protein
MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRIL
PDGTVDGTRDRSDQHIQLQLSAESVGEVYIKSTETGQYLAM
NLRGPYYIDGILKALNGGG
"""

# Regex pattern to match headers and their sequences
pattern = r">(.*?)\n([^>]+)"
matches = re.findall(pattern, fasta_data, re.DOTALL)

# Build the formatted FASTA string
formatted_fasta = ""
for header, seq in matches:
    # Remove newlines and extra whitespace from the sequence
    cleaned_seq = seq.replace("\n", "").strip()
    formatted_fasta += f">{header}\n{cleaned_seq}\n"

print(formatted_fasta)

How this works:

  • re.DOTALL makes the . character match newlines, so we can capture multi-line sequences without breaking.
  • ([^>]+) grabs all content until the next > (the start of a new header), ensuring we get the full sequence for each entry.
  • We then strip out all newlines from the captured sequence to condense it into a single line.
Solution 2: Line-by-Line Processing (Easier to Debug)

If regex feels like a black box, this approach iterates through each line explicitly—it's great for catching edge cases like empty lines or unexpected formatting quirks:

def format_fasta(input_data):
    formatted_output = []
    current_sequence = []
    
    for line in input_data.splitlines():
        line = line.strip()
        if not line:  # Skip empty lines to avoid extra spaces
            continue
        
        if line.startswith(">"):
            # If we were building a sequence, add it to the output first
            if current_sequence:
                formatted_output.append("".join(current_sequence))
                current_sequence = []
            # Add the new header line to the output
            formatted_output.append(line)
        else:
            # Accumulate sequence lines to join later
            current_sequence.append(line)
    
    # Don't forget to add the last sequence in the file!
    if current_sequence:
        formatted_output.append("".join(current_sequence))
    
    return "\n".join(formatted_output)

# Example usage
fasta_data = """>A/Nayarit/InDRE2/2009
MDSGKGSSQKSRPKTRAPRRPSEPAVKRPVVEEPPAAEESV
PPPEKRKPTTTASTQAETANPQVQPQPAAQPAPPVQPQAPQ
PQAPQPAPQPAPQPAPQPQPQPAPQPAPQPQPQPAPQPAPQ
PQPQPAPQPGAPQPQPQPPQPQPPQPQPQPQPQPQPQPQP
>Another_Protein
MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRIL
PDGTVDGTRDRSDQHIQLQLSAESVGEVYIKSTETGQYLAM
NLRGPYYIDGILKALNGGG
"""

print(format_fasta(fasta_data))

How this works:

  • We keep track of the current sequence being built as we loop through each line.
  • When we hit a header line (>), we finalize the previous sequence (if any) and start fresh for the new entry.
  • Sequence lines get added to a list, which we join into a single string once we hit the next header or reach the end of the file.

Reading from a File

If your FASTA is stored in a file instead of a string, just read it in first:

with open("your_proteins.fasta", "r") as file:
    fasta_data = file.read()

# Then use either solution above
formatted_result = format_fasta(fasta_data)

# To save the formatted output back to a new file:
with open("formatted_proteins.fasta", "w") as out_file:
    out_file.write(formatted_result)

内容的提问来源于stack exchange,提问作者Durian At

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:21:46