You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将FASTA蛋白质序列文件转换为指定结构的Pandas DataFrame

实现代码

针对90万条大体积FASTA文件的解析需求,优先使用内存友好、处理高效的逐行解析方案:

import pandas as pd
import re

# 暂存所有解析后的条目
records = []
current_entry = None

with open("sample.faa", "r", encoding="utf-8") as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        # 匹配新条目的注释行
        if line.startswith(">"):
            # 把已经读取完成的上一个条目存入列表
            if current_entry is not None:
                records.append(current_entry)
            # 正则拆分注释行的ID、蛋白名称、物种信息
            match_res = re.match(r"^(>\S+)\s+(.*?)\s+(\[.*\])$", line)
            if match_res:
                id_val, name_val, sapiens_val = match_res.groups()
                current_entry = {
                    "ID": id_val,
                    "name": name_val,
                    "sapiens": sapiens_val,
                    "sequence": ""
                }
        # 匹配序列行,拼接至当前条目的序列字段
        else:
            if current_entry is not None:
                current_entry["sequence"] += line
    # 读取结束后存入最后一个条目
    if current_entry is not None:
        records.append(current_entry)

# 一次性转换为DataFrame,处理90万条数据效率远高于逐行插入
df = pd.DataFrame(records)

执行完成后调用df.head()即可验证输出格式和你要求的预期结构完全一致。

原有代码的问题说明

  1. 没有识别FASTA文件>开头的注释行标识,用不存在的"X"字符匹配ID,逻辑完全错误
  2. 没有处理蛋白质序列多行拆分的情况,会把每一行序列当成独立条目
  3. 没有拆分注释行的三个字段,也缺失物种列的处理逻辑
  4. 用loc逐行插入数据,对于90万条规模的数据效率极低,会出现严重卡顿

内容的提问来源于stack exchange,提问作者Mohammad Alshehri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 20:36:03