如何用列表推导式过滤含指定禁用词的字符串列表?
Python过滤含任意禁用词字符串的高效实现
你原来的代码问题在于,ignore_paths not in x是判断整个禁用词列表是否作为一个子串出现在x里,这显然不符合需求——你要的是检查x里是否包含列表中任意一个禁用词。
基础正确实现
用any()函数配合列表推导式就能解决:
ignore_paths = ['.ipynb_checkpoints', 'New', '_calibration', 'images'] A = [x for x in lof1 if not any(word in x for word in ignore_paths)]
any(word in x for word in ignore_paths)会在x包含任意一个禁用词时返回True,取反后就筛选出了不包含任何禁用词的元素。
高效优化方案(适合15万条数据)
如果要处理大量数据,预编译正则表达式的速度会更优,因为正则匹配的底层是C实现,比Python层面的循环判断更快:
import re ignore_paths = ['.ipynb_checkpoints', 'New', '_calibration', 'images'] # 转义禁用词里的正则特殊字符,避免匹配出错 pattern = re.compile('|'.join(re.escape(word) for word in ignore_paths)) A = [x for x in lof1 if not pattern.search(x)]
这个方案在处理15万条路径时,性能会比纯Python循环的any()更出色。
示例验证
当输入:
lof1 = ['User/images/20210701_151111_G1100_E53100_r121_g64_b154_WBA0_GA0_EA0_6aa87af_crop.png', 'User/images/16f48a97-7111-4f66-92cc-dc7329e7ec92.png', 'User/images/image_2022_06_21-11_41_04_AM.png']
上述两种方案都会返回空列表,符合预期。
内容的提问来源于stack exchange,提问作者ali eskandari
相关产品推荐
相关产品推荐

