You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python按字符匹配两个序列文件并添加预测分数

解决方案

核心思路

要实现序列位置与预测分数的匹配,关键是定位结合区域在序列中的具体位置,再将分数映射到对应位置。pd.str.match()仅能判断整串匹配,无法定位子串位置,因此需要结合正则表达式提取子串的起始/结束索引,再逐个位置赋值分数。

假设文件格式

先明确两个文件的典型结构(可根据实际格式调整读取逻辑):

  • file1(序列文件):
    seq_id,sequence
    seq A,ABCDEFGHIJK
    seq B,LIMNOPQRSTU
    
  • file2(结合区域+分数文件):
    seq_id,binding_region,score
    seq A,CDE,0.85
    seq A,GHI,0.92
    seq B,MNO,0.78
    

具体实现代码

import pandas as pd
import re

# 1. 读取文件(根据实际格式调整参数,比如无表头用header=None)
df_seq = pd.read_csv("file1.csv")
df_binding = pd.read_csv("file2.csv")

# 2. 按seq_id合并两个数据集,保留所有序列
df_merged = pd.merge(df_seq, df_binding, on="seq_id", how="left")

# 3. 定义函数:将结合区域的分数映射到序列的每个位置
def map_score_to_positions(row):
    sequence = row["sequence"]
    binding_region = row["binding_region"]
    score = row["score"]
    
    # 初始化分数列表,未匹配位置默认设为0.0(可改为NaN)
    position_scores = [0.0] * len(sequence)
    
    # 处理空值(当序列无对应结合区域时)
    if pd.isna(binding_region):
        return pd.Series({
            "seq_id": row["seq_id"],
            "position": list(range(1, len(sequence)+1)),
            "amino_acid": list(sequence),
            "score": position_scores
        })
    
    # 找到结合区域在序列中的所有出现位置
    for match in re.finditer(re.escape(binding_region), sequence):
        start_idx = match.start()
        end_idx = match.end()
        # 给对应区间的位置赋值分数
        for i in range(start_idx, end_idx):
            position_scores[i] = score
    
    # 返回展开后的位置、氨基酸和分数
    return pd.Series({
        "seq_id": row["seq_id"],
        "position": list(range(1, len(sequence)+1)),
        "amino_acid": list(sequence),
        "score": position_scores
    })

# 4. 应用函数并展开成每行对应一个序列位置
result_df = df_merged.apply(map_score_to_positions, axis=1)
result_df = result_df.explode(["position", "amino_acid", "score"], ignore_index=True)

# 5. 处理重复位置(同一位置被多个结合区域覆盖时,保留最高分数)
result_df = result_df.groupby(["seq_id", "position", "amino_acid"])["score"].max().reset_index()

# 6. 保存结果表
result_df.to_csv("sequence_score_table.csv", index=False)

关键说明

  • 用re.finditer()精准定位结合区域在序列中的所有起始/结束索引,解决了pd.str.match()无法定位子串的问题;
  • 合并时使用how="left"确保所有序列都被保留,无结合区域的位置分数默认设为0.0(可根据需求改为pd.NA);
  • 通过explode()将每个序列的位置展开为单独行,便于后续分析;
  • 新增groupby().max()处理同一位置被多个结合区域覆盖的场景,保留最高预测分数。

内容的提问来源于stack exchange,提问作者worfei alafei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 04:55:32