You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python(Google Colab)中将获取的核酸序列导出为FASTA格式

实现方案

你现有的检索逻辑不需要大幅调整,补充文件写入和Colab下载逻辑即可,适配批量序列场景,输出格式完全匹配你给出的示例要求。

注意事项

  • 必须先配置Entrez.email为你的个人邮箱,这是NCBI对Entrez接口的强制要求,批量请求时不填很容易被拦截导致请求失败。

完整修改后代码

# 导入所需模块
from Bio import SeqIO
from Bio import Entrez
from google.colab import files

# 配置NCBI请求邮箱,替换为你自己的邮箱
Entrez.email = "your_email@example.com"

# 读取检索ID
lines = []
with open('my_file.seq') as f:
  lines = [line.rstrip() for line in f] 

# 检索序列并写入文件
output_file = "seq_output.txt"
with open(output_file, "w") as f_out:
  handle = Entrez.efetch(db="nucleotide", id=lines, rettype="gb", retmode="text")
  record_iterator = SeqIO.parse(handle, "gb")
  for record in record_iterator:
    # 按ID+空格+序列的格式写入,每行一条
    f_out.write(f"{record.id} {record.seq}\n")
  handle.close()

# 触发Colab文件下载
files.download(output_file)

额外说明

如果后续你需要标准FASTA格式(>开头带序列描述的版本),可以直接用Biopython内置的SeqIO.write方法,不需要手动拼接内容,批量处理效率更高:

# 标准FASTA格式导出示例
with open("output.fasta", "w") as f_out:
  SeqIO.write(record_iterator, f_out, "fasta")

内容的提问来源于stack exchange,提问作者Camp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 07:36:03