如何用Python按字符匹配两个序列文件并添加预测分数
解决方案
核心思路
要实现序列位置与预测分数的匹配,关键是定位结合区域在序列中的具体位置,再将分数映射到对应位置。pd.str.match()仅能判断整串匹配,无法定位子串位置,因此需要结合正则表达式提取子串的起始/结束索引,再逐个位置赋值分数。
假设文件格式
先明确两个文件的典型结构(可根据实际格式调整读取逻辑):
- file1(序列文件):
seq_id,sequence seq A,ABCDEFGHIJK seq B,LIMNOPQRSTU - file2(结合区域+分数文件):
seq_id,binding_region,score seq A,CDE,0.85 seq A,GHI,0.92 seq B,MNO,0.78
具体实现代码
import pandas as pd import re # 1. 读取文件(根据实际格式调整参数,比如无表头用header=None) df_seq = pd.read_csv("file1.csv") df_binding = pd.read_csv("file2.csv") # 2. 按seq_id合并两个数据集,保留所有序列 df_merged = pd.merge(df_seq, df_binding, on="seq_id", how="left") # 3. 定义函数:将结合区域的分数映射到序列的每个位置 def map_score_to_positions(row): sequence = row["sequence"] binding_region = row["binding_region"] score = row["score"] # 初始化分数列表,未匹配位置默认设为0.0(可改为NaN) position_scores = [0.0] * len(sequence) # 处理空值(当序列无对应结合区域时) if pd.isna(binding_region): return pd.Series({ "seq_id": row["seq_id"], "position": list(range(1, len(sequence)+1)), "amino_acid": list(sequence), "score": position_scores }) # 找到结合区域在序列中的所有出现位置 for match in re.finditer(re.escape(binding_region), sequence): start_idx = match.start() end_idx = match.end() # 给对应区间的位置赋值分数 for i in range(start_idx, end_idx): position_scores[i] = score # 返回展开后的位置、氨基酸和分数 return pd.Series({ "seq_id": row["seq_id"], "position": list(range(1, len(sequence)+1)), "amino_acid": list(sequence), "score": position_scores }) # 4. 应用函数并展开成每行对应一个序列位置 result_df = df_merged.apply(map_score_to_positions, axis=1) result_df = result_df.explode(["position", "amino_acid", "score"], ignore_index=True) # 5. 处理重复位置(同一位置被多个结合区域覆盖时,保留最高分数) result_df = result_df.groupby(["seq_id", "position", "amino_acid"])["score"].max().reset_index() # 6. 保存结果表 result_df.to_csv("sequence_score_table.csv", index=False)
关键说明
- 用
re.finditer()精准定位结合区域在序列中的所有起始/结束索引,解决了pd.str.match()无法定位子串的问题; - 合并时使用
how="left"确保所有序列都被保留,无结合区域的位置分数默认设为0.0(可根据需求改为pd.NA); - 通过
explode()将每个序列的位置展开为单独行,便于后续分析; - 新增
groupby().max()处理同一位置被多个结合区域覆盖的场景,保留最高预测分数。
内容的提问来源于stack exchange,提问作者worfei alafei
相关产品推荐
相关产品推荐

