如何将FASTA蛋白质序列文件转换为指定结构的Pandas DataFrame
实现代码
针对90万条大体积FASTA文件的解析需求,优先使用内存友好、处理高效的逐行解析方案:
import pandas as pd import re # 暂存所有解析后的条目 records = [] current_entry = None with open("sample.faa", "r", encoding="utf-8") as f: for line in f: line = line.strip() if not line: continue # 匹配新条目的注释行 if line.startswith(">"): # 把已经读取完成的上一个条目存入列表 if current_entry is not None: records.append(current_entry) # 正则拆分注释行的ID、蛋白名称、物种信息 match_res = re.match(r"^(>\S+)\s+(.*?)\s+(\[.*\])$", line) if match_res: id_val, name_val, sapiens_val = match_res.groups() current_entry = { "ID": id_val, "name": name_val, "sapiens": sapiens_val, "sequence": "" } # 匹配序列行,拼接至当前条目的序列字段 else: if current_entry is not None: current_entry["sequence"] += line # 读取结束后存入最后一个条目 if current_entry is not None: records.append(current_entry) # 一次性转换为DataFrame,处理90万条数据效率远高于逐行插入 df = pd.DataFrame(records)
执行完成后调用df.head()即可验证输出格式和你要求的预期结构完全一致。
原有代码的问题说明
- 没有识别FASTA文件
>开头的注释行标识,用不存在的"X"字符匹配ID,逻辑完全错误 - 没有处理蛋白质序列多行拆分的情况,会把每一行序列当成独立条目
- 没有拆分注释行的三个字段,也缺失物种列的处理逻辑
- 用
loc逐行插入数据,对于90万条规模的数据效率极低,会出现严重卡顿
内容的提问来源于stack exchange,提问作者Mohammad Alshehri
相关产品推荐
相关产品推荐

