大小写不敏感筛选文件路径时如何获取匹配项的原始索引
原有代码问题
- for循环后直接跟列表推导属于语法错误,无法正常运行
- index()方法仅能返回第一个匹配项的位置,遇到重复路径会返回错误索引,原有返回逻辑完全不匹配需求
实现方案
你可以通过enumerate遍历小写路径列表同时记录原始索引,配合后缀匹配判断即可得到目标索引列表,以下是可直接运行的实现:
如果你需要返回需保留条目对应的原始索引,使用以下代码:
def index_pos_filtered_paths(path_list_lowercase, suffixes_unwanted): # 统一将不需要的后缀转小写,避免大小写漏匹配 suffixes_unwanted_lower = [s.lower() for s in suffixes_unwanted] keep_indexes = [] for idx, path_lower in enumerate(path_list_lowercase): # 路径不包含任何不需要的后缀时,记录当前索引 if all(suffix not in path_lower for suffix in suffixes_unwanted_lower): keep_indexes.append(idx) return keep_indexes index_of_filtered_paths_list = index_pos_filtered_paths(path_list_lowercase, suffixes_unwanted)
如果你需要返回需删除条目对应的原始索引,仅需调整判断条件即可:
def index_pos_filtered_paths(path_list_lowercase, suffixes_unwanted): suffixes_unwanted_lower = [s.lower() for s in suffixes_unwanted] del_indexes = [] for idx, path_lower in enumerate(path_list_lowercase): # 路径命中任意一个不需要的后缀时,记录当前索引 if any(suffix in path_lower for suffix in suffixes_unwanted_lower): del_indexes.append(idx) return del_indexes index_of_filtered_paths_list = index_pos_filtered_paths(path_list_lowercase, suffixes_unwanted)
后续操作注意事项
- 操作Pandas DataFrame时,直接调用
df.drop(index_of_filtered_paths_list, inplace=True)即可完成对应行的删除,完全不会改动原始路径的大小写格式。 - 操作Python原生列表时,注意从大索引到小索引的顺序删除,避免删除前置元素后导致后续索引偏移出错。
内容的提问来源于stack exchange,提问作者cowboykevin05
相关产品推荐
相关产品推荐

