Python中根据匹配标记数组筛选图书标题数组的方法
根据匹配标记过滤标题数组
在爬取图书ISBN时,搜索书名常出现非目标结果。比如搜索"The third bear Jeff VanderMeer"返回的结果里,存在无关条目,需要根据预先生成的匹配标记数组过滤标题,保留目标内容。
现有输入数据:
aux_title = [ " The Year's Top Ten Tales of Science Fiction", ' The Third Bear', ' The Third Bear', ' The Third Bear' ] match_names = [0., 1., 1., 1.]
需要过滤后得到仅包含目标标题的数组:
[' The Third Bear', ' The Third Bear', ' The Third Bear']
几种简便实现方法(类似Matlab find 功能)
1. 列表推导式(简单直观,适合中小数据量)
直接遍历标题和标记的配对,筛选标记为1的条目:
filtered_titles = [title for title, match in zip(aux_title, match_names) if match == 1.]
2. Numpy 索引(高效处理大数据量,适配22000条数据)
如果数据量较大,用numpy的布尔索引会更高效,逻辑和Matlab的find用法一致:
import numpy as np aux_title_np = np.array(aux_title) match_names_np = np.array(match_names) filtered_titles = aux_title_np[match_names_np == 1.].tolist()
numpy的向量化操作在处理大规模数据时性能优于纯Python循环,22000条数据完全无压力。
3. 使用filter函数(函数式风格)
结合lambda表达式完成过滤:
# 先配对标题和标记,过滤出匹配项 filtered_pairs = list(filter(lambda x: x[1] == 1., zip(aux_title, match_names))) # 提取标题部分 filtered_titles = [item[0] for item in filtered_pairs]
以上方法都能快速完成过滤需求,其中numpy方法在处理大数据时优势明显,列表推导式则更简洁易读,可根据实际场景选择。
内容的提问来源于stack exchange,提问作者Maximiliano Machado Goncalves
相关产品推荐
相关产品推荐

