Python中获取蛋白质基序在序列中的起止位置方法咨询
获取蛋白质序列中基序的起始与终止位置
你可以通过自定义函数结合re.finditer()提取基序在序列中的起始和终止位置,以下是具体实现方案:
方法1:提取所有匹配的位置(支持多基序场景)
先定义处理单个序列的函数,返回所有匹配基序的位置信息:
import re import pandas as pd def get_motif_positions(sequence, motif): positions = [] # 遍历序列中所有符合的基序匹配 for match in re.finditer(motif, sequence): # start()是起始索引,end()是终止索引(Python字符串左闭右开,end = start + 基序长度) positions.append((match.start(), match.end())) return positions
将函数应用到你筛选后的数据集:
# 先筛选包含目标基序的行 filtered_df = df.loc[df['Protein_sequence'].str.contains("WA[T]R", regex=True)].copy() # 新增列存储基序位置 filtered_df['Motif_positions'] = filtered_df['Protein_sequence'].apply(lambda x: get_motif_positions(x, "WA[T]R"))
执行后filtered_df的Motif_positions列会存储每个序列中所有匹配基序的位置元组,例如[(12, 16)](对应基序从第12位开始,第16位结束)。
方法2:仅提取第一个匹配的位置
如果你的数据集里每个序列仅存在一个匹配基序,可以简化代码:
def get_first_motif_position(sequence, motif): match = re.search(motif, sequence) if match: return (match.start(), match.end()) return None filtered_df['First_motif_position'] = filtered_df['Protein_sequence'].apply(lambda x: get_first_motif_position(x, "WA[T]R"))
额外说明
- Python字符串索引从0开始,若需要转换成生物学常用的1起始索引,只需将起始位置+1,比如
(match.start()+1, match.end())。
内容的提问来源于stack exchange,提问作者Rajdeep Jaswal
相关产品推荐
相关产品推荐

