如何去除Python列表多个元素中重复的“anti-human ”前缀文本
原有代码问题说明
Python中字符串是不可变类型,调用split()方法只会返回新的分割结果,不会修改原字符串本身,且你遍历过程中没有存储处理后的结果,因此无法得到预期输出。
推荐解决方案
方案1:使用str.removeprefix()(Python 3.9及以上版本,最稳妥)
该方法专门用于移除字符串开头的指定前缀,不会误处理字符串中间出现的相同内容,完美匹配你的需求:
# parsed_protein_names 为你的原始抗体名称列表 target_proteins = [item.removeprefix('anti-human ') for item in parsed_protein_names]
运行后即可得到目标结果,比如'anti-human CD274 (B7-H1, PD-L1)'会被处理为'CD274 (B7-H1, PD-L1)'。
方案2:兼容低版本Python的实现
如果你的Python版本低于3.9,没有removeprefix方法,可以用split结合索引的方式实现,注意要取分割后的第二个元素并存储结果:
target_proteins = [] for item in parsed_protein_names: # 按前缀分割后,第一个元素是空字符串,第二个元素就是你需要的蛋白靶点内容 processed_item = item.split('anti-human ')[1] target_proteins.append(processed_item)
也可以用列表推导式简化写法:
target_proteins = [item.split('anti-human ')[1] for item in parsed_protein_names]
如果你的列表中可能存在不带anti-human 前缀的异常元素,可以加判断避免索引报错:
target_proteins = [ item.split('anti-human ')[1] if item.startswith('anti-human ') else item for item in parsed_protein_names ]
内容的提问来源于stack exchange,提问作者Michael Mather
相关产品推荐
相关产品推荐

