You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在DataFrame列字符串中查找目标词并输出词位置及对应Record ID

高效搜索DataFrame文本列中目标词的词位置方案

针对大型数据集,优先采用矢量化操作和正则工具实现高效搜索,以下是满足需求的具体方案:

实现步骤

1. 示例数据准备

先构造模拟数据还原场景:

import pandas as pd
import re

data = {
    "Record ID": [1, 2, 3],
    "Text": [
        "hello world hello python",
        "python is great, python is fun",
        "data science with python"
    ]
}
df = pd.DataFrame(data)

2. 核心搜索逻辑

定义函数提取目标词的词位置,再通过矢量化方法批量处理:

def get_target_positions(text, target):
    # 按空白分割文本为词列表(可根据需求调整分词规则,比如过滤标点)
    words = re.split(r'\s+', text.strip())
    # 收集所有匹配目标词的位置(从1开始计数)
    return [idx + 1 for idx, word in enumerate(words) if word.lower() == target.lower()]

# 指定目标词,批量处理Text列
target_word = "python"
df["Word Positions"] = df["Text"].apply(lambda x: get_target_positions(x, target_word))

# 展开多值结果,保留Record ID
result_df = df.explode("Word Positions").dropna(subset=["Word Positions"])

3. 大数据场景优化

如果处理百万级以上数据,可通过以下方式进一步提升效率:

  • 使用swifter库自动适配并行/矢量化执行,避免手动优化:
import swifter

df["Word Positions"] = df["Text"].swifter.apply(lambda x: get_target_positions(x, target_word))
  • 若需更精准的分词(如处理带标点的词汇),可提前加载nltk分词器,避免重复初始化开销。

输出结果

最终result_df的结构如下:

Record IDTextWord Positions
1hello world hello python4
2python is great, python is fun1
2python is great, python is fun4
3data science with python4

关键注意事项

  • 词位置计算以空白分割后的词索引+1为准,若需自定义分词规则(如去除标点),可修改re.split的正则表达式。
  • 代码默认大小写不敏感匹配,若需精确匹配,移除lower()转换即可。
  • 矢量化操作的效率远高于逐行循环,适合处理从数据库导入的大型DataFrame。

内容的提问来源于stack exchange,提问作者tshobe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 22:35:51