You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅对FASTA文件中DNA序列翻译并保留元数据格式

解决FASTA序列按条目翻译并保留元数据的问题

核心思路

要实现每个FASTA条目单独翻译、保留元数据,关键是正确解析FASTA格式:把>开头的行作为元数据标记,对应收集后续所有非>行的DNA序列,完成一个条目后立即处理翻译,再处理下一个条目,避免跨条目合并序列。

代码实现(Python)

以下是可直接运行的代码,包含FASTA解析、DNA→RNA→蛋白质的完整流程:

# 标准密码子表(RNA→氨基酸)
codon_table = {
    'UUU': 'F', 'UUC': 'F', 'UUA': 'L', 'UUG': 'L',
    'UCU': 'S', 'UCC': 'S', 'UCA': 'S', 'UCG': 'S',
    'UAU': 'Y', 'UAC': 'Y', 'UAA': '*', 'UAG': '*',
    'UGU': 'C', 'UGC': 'C', 'UGA': '*', 'UGG': 'W',
    'CUU': 'L', 'CUC': 'L', 'CUA': 'L', 'CUG': 'L',
    'CCU': 'P', 'CCC': 'P', 'CCA': 'P', 'CCG': 'P',
    'CAU': 'H', 'CAC': 'H', 'CAA': 'Q', 'CAG': 'Q',
    'CGU': 'R', 'CGC': 'R', 'CGA': 'R', 'CGG': 'R',
    'AUU': 'I', 'AUC': 'I', 'AUA': 'I', 'AUG': 'M',
    'ACU': 'T', 'ACC': 'T', 'ACA': 'T', 'ACG': 'T',
    'AAU': 'N', 'AAC': 'N', 'AAA': 'K', 'AAG': 'K',
    'AGU': 'S', 'AGC': 'S', 'AGA': 'R', 'AGG': 'R',
    'GUU': 'V', 'GUC': 'V', 'GUA': 'V', 'GUG': 'V',
    'GCU': 'A', 'GCC': 'A', 'GCA': 'A', 'GCG': 'A',
    'GAU': 'D', 'GAC': 'D', 'GAA': 'E', 'GAG': 'E',
    'GGU': 'G', 'GGC': 'G', 'GGA': 'G', 'GGG': 'G'
}

def dna_to_rna(dna_seq):
    # DNA转RNA:替换T为U,兼容大小写输入
    return dna_seq.replace('T', 'U').replace('t', 'u')

def rna_to_protein(rna_seq):
    # RNA转蛋白质:按3个碱基一组翻译,遇到终止密码子停止
    protein = []
    for i in range(0, len(rna_seq) - 2, 3):
        codon = rna_seq[i:i+3]
        amino_acid = codon_table.get(codon.upper(), '?')  # 未知密码子用?代替
        if amino_acid == '*':
            break  # 遇到终止密码子停止翻译,需全序列翻译可删除此行
        protein.append(amino_acid)
    return ''.join(protein)

def process_fasta(input_file, output_file):
    current_header = None
    current_dna = []
    
    with open(input_file, 'r') as infile, open(output_file, 'w') as outfile:
        for line in infile:
            line = line.strip()
            if not line:
                continue  # 跳过空行
            if line.startswith('>'):
                # 处理上一个条目(如果存在)
                if current_header and current_dna:
                    dna_seq = ''.join(current_dna)
                    rna_seq = dna_to_rna(dna_seq)
                    protein_seq = rna_to_protein(rna_seq)
                    # 按FASTA格式写入结果
                    outfile.write(f"{current_header}\n")
                    outfile.write(f"{protein_seq}\n")
                # 更新当前条目的元数据和序列容器
                current_header = line
                current_dna = []
            else:
                # 收集当前条目的DNA序列(统一转为大写)
                current_dna.append(line.upper())
        # 处理最后一个条目
        if current_header and current_dna:
            dna_seq = ''.join(current_dna)
            rna_seq = dna_to_rna(dna_seq)
            protein_seq = rna_to_protein(rna_seq)
            outfile.write(f"{current_header}\n")
            outfile.write(f"{protein_seq}\n")

# 使用示例:替换为你的输入输出文件路径
process_fasta("input.fasta", "output_proteins.fasta")

代码说明

  • FASTA解析逻辑:逐行遍历文件,遇到>行就触发上一个条目的翻译流程,同时记录新的元数据;非>行拼接为当前条目的DNA序列,确保不同条目完全独立处理。
  • 序列转换:
    • DNA转RNA直接替换T为U,兼容大小写输入;
    • RNA转蛋白质使用标准密码子表,遇到终止密码子*停止翻译,未知密码子用?标记。
  • 输出格式:严格遵循FASTA规范,每个>元数据对应下方一行翻译后的蛋白序列,与输入结构完全匹配。

自定义调整

  • 若需要翻译整个序列(不提前终止),删除if amino_acid == '*': break即可;
  • 若要过滤非ATCG字符,可在收集DNA序列时添加line = ''.join([c for c in line if c.upper() in 'ATCG']);
  • 支持多段换行的DNA序列自动拼接,无需额外处理。

内容的提问来源于stack exchange,提问作者ooCrafty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 03:01:39