You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中获取蛋白质基序在序列中的起止位置方法咨询

获取蛋白质序列中基序的起始与终止位置

你可以通过自定义函数结合re.finditer()提取基序在序列中的起始和终止位置,以下是具体实现方案:

方法1:提取所有匹配的位置(支持多基序场景)

先定义处理单个序列的函数,返回所有匹配基序的位置信息:

import re
import pandas as pd

def get_motif_positions(sequence, motif):
    positions = []
    # 遍历序列中所有符合的基序匹配
    for match in re.finditer(motif, sequence):
        # start()是起始索引,end()是终止索引(Python字符串左闭右开,end = start + 基序长度)
        positions.append((match.start(), match.end()))
    return positions

将函数应用到你筛选后的数据集:

# 先筛选包含目标基序的行
filtered_df = df.loc[df['Protein_sequence'].str.contains("WA[T]R", regex=True)].copy()

# 新增列存储基序位置
filtered_df['Motif_positions'] = filtered_df['Protein_sequence'].apply(lambda x: get_motif_positions(x, "WA[T]R"))

执行后filtered_df的Motif_positions列会存储每个序列中所有匹配基序的位置元组,例如[(12, 16)](对应基序从第12位开始,第16位结束)。

方法2:仅提取第一个匹配的位置

如果你的数据集里每个序列仅存在一个匹配基序,可以简化代码:

def get_first_motif_position(sequence, motif):
    match = re.search(motif, sequence)
    if match:
        return (match.start(), match.end())
    return None

filtered_df['First_motif_position'] = filtered_df['Protein_sequence'].apply(lambda x: get_first_motif_position(x, "WA[T]R"))

额外说明

  • Python字符串索引从0开始,若需要转换成生物学常用的1起始索引,只需将起始位置+1,比如(match.start()+1, match.end())。

内容的提问来源于stack exchange,提问作者Rajdeep Jaswal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 20:35:18