You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas两个字符串Series逐行匹配并保留长Series原索引的实现问题

实现方案

直接使用pandas的apply方法逐行处理文本Series,天然保证输出结果与原文本Series的索引完全对齐,不会出现索引错位问题。

核心逻辑

  1. 先将标签Series转为列表,避免重复遍历Series对象
  2. 对文本Series的每行内容,筛选出所有存在于标签列表中的字符串
  3. apply方法会自动保留原Series的索引,无需手动维护索引对应关系

完整代码

import pandas as pd

# 1. 将标签转为列表
label_list = labels['uniq_labels'].tolist()

# 2. 逐行匹配,输出结果索引与reports['corpus']完全一致
matched_series = reports['corpus'].apply(lambda x: [label for label in label_list if label in x])

示例验证

用你提供的示例数据运行代码:

reports_series = pd.Series(['This is a test first sentence. \
                            This is the first row of a pandas series.',
                            'Here is the second row. The row that means the most. The row that never goes away.',
                            'The third sentence. The third row to the example pandas series.',
                            'This is the fourth and only fourth row of the pandas series.',
                            'Here is the fifth row. The fifth row that means the most.'])

labels_series = pd.Series(['first', 'sentence', 'second row'])

label_list = labels_series.tolist()
matched_series = reports_series.apply(lambda x: [label for label in label_list if label in x])

输出的matched_series结果如下,索引与原reports_series完全对应:

0    [first, sentence]
1        [second row]
2          [sentence]
3                   []
4                   []
dtype: object

可选优化:整词精确匹配

如果需要避免子串误匹配(比如标签为cat时,不会把文本中的category判定为匹配),可以用正则词边界实现严格整词匹配:

import re

label_list = labels['uniq_labels'].tolist()
# 增加词边界校验,re.escape用于处理标签中含特殊正则字符的情况
matched_series = reports['corpus'].apply(lambda x: [label for label in label_list if re.search(rf'\b{re.escape(label)}\b', x) is not None])

内容的提问来源于stack exchange,提问作者eliboy8

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:39:03