You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中根据匹配标记数组筛选图书标题数组的方法

根据匹配标记过滤标题数组

在爬取图书ISBN时,搜索书名常出现非目标结果。比如搜索"The third bear Jeff VanderMeer"返回的结果里,存在无关条目,需要根据预先生成的匹配标记数组过滤标题,保留目标内容。

现有输入数据:

aux_title = [
    " The Year's Top Ten Tales of Science Fiction",
    ' The Third Bear',
    ' The Third Bear',
    ' The Third Bear'
]
match_names = [0., 1., 1., 1.]

需要过滤后得到仅包含目标标题的数组:

[' The Third Bear', ' The Third Bear', ' The Third Bear']

几种简便实现方法(类似Matlab find 功能)

1. 列表推导式(简单直观,适合中小数据量)

直接遍历标题和标记的配对,筛选标记为1的条目:

filtered_titles = [title for title, match in zip(aux_title, match_names) if match == 1.]

2. Numpy 索引(高效处理大数据量,适配22000条数据)

如果数据量较大,用numpy的布尔索引会更高效,逻辑和Matlab的find用法一致:

import numpy as np

aux_title_np = np.array(aux_title)
match_names_np = np.array(match_names)
filtered_titles = aux_title_np[match_names_np == 1.].tolist()

numpy的向量化操作在处理大规模数据时性能优于纯Python循环,22000条数据完全无压力。

3. 使用filter函数(函数式风格)

结合lambda表达式完成过滤:

# 先配对标题和标记,过滤出匹配项
filtered_pairs = list(filter(lambda x: x[1] == 1., zip(aux_title, match_names)))
# 提取标题部分
filtered_titles = [item[0] for item in filtered_pairs]

以上方法都能快速完成过滤需求,其中numpy方法在处理大数据时优势明显,列表推导式则更简洁易读,可根据实际场景选择。

内容的提问来源于stack exchange,提问作者Maximiliano Machado Goncalves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 19:12:40