如何仅对FASTA文件中DNA序列翻译并保留元数据格式
解决FASTA序列按条目翻译并保留元数据的问题
核心思路
要实现每个FASTA条目单独翻译、保留元数据,关键是正确解析FASTA格式:把>开头的行作为元数据标记,对应收集后续所有非>行的DNA序列,完成一个条目后立即处理翻译,再处理下一个条目,避免跨条目合并序列。
代码实现(Python)
以下是可直接运行的代码,包含FASTA解析、DNA→RNA→蛋白质的完整流程:
# 标准密码子表(RNA→氨基酸) codon_table = { 'UUU': 'F', 'UUC': 'F', 'UUA': 'L', 'UUG': 'L', 'UCU': 'S', 'UCC': 'S', 'UCA': 'S', 'UCG': 'S', 'UAU': 'Y', 'UAC': 'Y', 'UAA': '*', 'UAG': '*', 'UGU': 'C', 'UGC': 'C', 'UGA': '*', 'UGG': 'W', 'CUU': 'L', 'CUC': 'L', 'CUA': 'L', 'CUG': 'L', 'CCU': 'P', 'CCC': 'P', 'CCA': 'P', 'CCG': 'P', 'CAU': 'H', 'CAC': 'H', 'CAA': 'Q', 'CAG': 'Q', 'CGU': 'R', 'CGC': 'R', 'CGA': 'R', 'CGG': 'R', 'AUU': 'I', 'AUC': 'I', 'AUA': 'I', 'AUG': 'M', 'ACU': 'T', 'ACC': 'T', 'ACA': 'T', 'ACG': 'T', 'AAU': 'N', 'AAC': 'N', 'AAA': 'K', 'AAG': 'K', 'AGU': 'S', 'AGC': 'S', 'AGA': 'R', 'AGG': 'R', 'GUU': 'V', 'GUC': 'V', 'GUA': 'V', 'GUG': 'V', 'GCU': 'A', 'GCC': 'A', 'GCA': 'A', 'GCG': 'A', 'GAU': 'D', 'GAC': 'D', 'GAA': 'E', 'GAG': 'E', 'GGU': 'G', 'GGC': 'G', 'GGA': 'G', 'GGG': 'G' } def dna_to_rna(dna_seq): # DNA转RNA:替换T为U,兼容大小写输入 return dna_seq.replace('T', 'U').replace('t', 'u') def rna_to_protein(rna_seq): # RNA转蛋白质:按3个碱基一组翻译,遇到终止密码子停止 protein = [] for i in range(0, len(rna_seq) - 2, 3): codon = rna_seq[i:i+3] amino_acid = codon_table.get(codon.upper(), '?') # 未知密码子用?代替 if amino_acid == '*': break # 遇到终止密码子停止翻译,需全序列翻译可删除此行 protein.append(amino_acid) return ''.join(protein) def process_fasta(input_file, output_file): current_header = None current_dna = [] with open(input_file, 'r') as infile, open(output_file, 'w') as outfile: for line in infile: line = line.strip() if not line: continue # 跳过空行 if line.startswith('>'): # 处理上一个条目(如果存在) if current_header and current_dna: dna_seq = ''.join(current_dna) rna_seq = dna_to_rna(dna_seq) protein_seq = rna_to_protein(rna_seq) # 按FASTA格式写入结果 outfile.write(f"{current_header}\n") outfile.write(f"{protein_seq}\n") # 更新当前条目的元数据和序列容器 current_header = line current_dna = [] else: # 收集当前条目的DNA序列(统一转为大写) current_dna.append(line.upper()) # 处理最后一个条目 if current_header and current_dna: dna_seq = ''.join(current_dna) rna_seq = dna_to_rna(dna_seq) protein_seq = rna_to_protein(rna_seq) outfile.write(f"{current_header}\n") outfile.write(f"{protein_seq}\n") # 使用示例:替换为你的输入输出文件路径 process_fasta("input.fasta", "output_proteins.fasta")
代码说明
- FASTA解析逻辑:逐行遍历文件,遇到
>行就触发上一个条目的翻译流程,同时记录新的元数据;非>行拼接为当前条目的DNA序列,确保不同条目完全独立处理。 - 序列转换:
- DNA转RNA直接替换
T为U,兼容大小写输入; - RNA转蛋白质使用标准密码子表,遇到终止密码子
*停止翻译,未知密码子用?标记。
- DNA转RNA直接替换
- 输出格式:严格遵循FASTA规范,每个
>元数据对应下方一行翻译后的蛋白序列,与输入结构完全匹配。
自定义调整
- 若需要翻译整个序列(不提前终止),删除
if amino_acid == '*': break即可; - 若要过滤非ATCG字符,可在收集DNA序列时添加
line = ''.join([c for c in line if c.upper() in 'ATCG']); - 支持多段换行的DNA序列自动拼接,无需额外处理。
内容的提问来源于stack exchange,提问作者ooCrafty
相关产品推荐
相关产品推荐

