Python中如何通过正则匹配简洁移除列表中多个不需要的元素
Excel工作表筛选逻辑精简实现
需求说明
现有Excel工作簿中包含两类工作表:
- 一类以国家名称命名,为需要保留的目标工作表
- 其余为无关工作表,需要全部过滤剔除
原有实现代码可正常输出预期结果:['India', 'Brazil', 'Czech', 'Italy', 'Hungary', 'Poland'],但代码冗余度较高:分4条独立语句收集待移除工作表名称,存在大量重复逻辑,需要更简洁的实现方案。
原有冗余实现代码如下:
all_sheets = xl.sheet_names # see all sheet names entries_to_remove_1 = ['Information', 'Summary', 'Template', 'Source', 'BOW'] entries_to_remove_2 = list(filter(lambda x: 'Annual' in x, all_sheets)) entries_to_remove_3 = list(filter(lambda x: 'Quarterly' in x, all_sheets)) entries_to_remove_4 = list(filter(lambda x: 'WIP' in x, all_sheets)) entries_to_remove = entries_to_remove_1 + entries_to_remove_2 + entries_to_remove_3 + entries_to_remove_4 entries_for_replacement = [i for i in all_sheets if i not in entries_to_remove] all_sheets = list(set(entries_for_replacement)) all_sheets
精简优化方案
核心优化思路是整合所有排除规则,去掉冗余的中间待移除列表,直接在筛选阶段完成判断,减少不必要的变量定义:
all_sheets = xl.sheet_names # 定义两类排除规则:精确匹配的表名、需要模糊匹配的关键词 exclude_exact = {'Information', 'Summary', 'Template', 'Source', 'BOW'} exclude_keywords = {'Annual', 'Quarterly', 'WIP'} # 集合推导一步完成过滤+去重 country_sheets = list({ sheet for sheet in all_sheets if sheet not in exclude_exact and not any(k in sheet for k in exclude_keywords) })
优化点说明
- 用集合存储排除项,成员判断的时间复杂度更低,数据量大时运行效率更高
- 所有模糊匹配规则通过
any()统一判断,不需要为每个关键词单独写filter逻辑 - 直接通过集合推导完成过滤和去重,省去了先生成待移除列表、再二次筛选的冗余步骤
可选保序版本
原实现用set()去重会打乱工作表原有顺序,如果需要保留工作表在Excel文件中的原始排列顺序(Python 3.7+ 字典默认保序),可以用如下写法:
country_sheets = list(dict.fromkeys([ sheet for sheet in all_sheets if sheet not in exclude_exact and not any(k in sheet for k in exclude_keywords) ]))
内容的提问来源于stack exchange,提问作者RSM
相关产品推荐
相关产品推荐

