如何计算DataFrame每行中多个突变位点的间距?
提取突变位点并计算相邻位点距离
完整解决方案代码
import pandas as pd import re # 示例数据 data = [[31581, "wild-type"], [31614, "D250C,E296C"], [31731, "T112K,T116I,E324I,S150C,N157C,V189C,D332C"]] df = pd.DataFrame(data, columns=['SAMPLE', 'MUTATION']) # 1. 过滤野生型样本,清理突变字段的冗余说明 df_clean = df[df['MUTATION'] != 'wild-type'].copy() # 移除所有"(Based on ...)"格式的后缀文本 df_clean['MUTATION'] = df_clean['MUTATION'].str.replace(r'\s*\(Based on .+\)', '', regex=True) # 2. 定义函数:提取突变位置并计算相邻位点距离 def get_mutation_distances(mutation_str): # 提取所有突变位置的数字 position_strs = re.findall(r'\d+', mutation_str) # 转换为整数列表 positions = list(map(int, position_strs)) # 按位置从大到小排序 sorted_positions = sorted(positions, reverse=True) # 计算相邻位点的距离(大位置减小位置) distances = [sorted_positions[i] - sorted_positions[i+1] for i in range(len(sorted_positions)-1)] return distances # 3. 批量处理所有样本,生成距离列表 df_clean['MUTATION_DISTANCES'] = df_clean['MUTATION'].apply(get_mutation_distances) # 查看结果 print(df_clean)
代码分步说明
数据过滤与清理
- 先排除标记为
wild-type的样本行,用copy()避免pandas的修改警告。 - 用正则表达式替换掉所有
(Based on ...)格式的冗余说明文本,比strip()更可靠(strip()仅处理字符串首尾字符,正则能匹配任意位置的这类文本)。
- 先排除标记为
提取突变位置
- 用
re.findall(r'\d+', mutation_str)直接从突变字符串中提取所有数字(即突变位点),无需拆分字符串后再过滤冗余字母,效率更高。
- 用
计算相邻位点距离
- 将提取的数字转为整数,按降序排序(确保后续计算的是大位点减小位点)。
- 通过列表推导式遍历排序后的位点,计算相邻两个位点的差值,生成距离列表。
批量处理
- 用
apply()函数对每行的MUTATION字段应用自定义函数,适合处理3万行以上的大型数据集,比手动循环更高效。
- 用
示例输出
SAMPLE MUTATION MUTATION_DISTANCES 1 31614 D250C,E296C [46] 2 31731 T112K,T116I,E324I,S150C,N157C,V189C,D332C [8, 135, 32, 7, 34, 4]
可选优化
- 如果需要去除重复的突变位点,可在转换整数后添加去重步骤:
positions = list(set(map(int, position_strs))) - 若希望按位点从小到大排序后计算差值(后位点减前位点),只需将
sorted(positions, reverse=True)改为sorted(positions)。
内容的提问来源于stack exchange,提问作者Paulina Brown
相关产品推荐
相关产品推荐

