如何在DataFrame列字符串中查找目标词并输出词位置及对应Record ID
高效搜索DataFrame文本列中目标词的词位置方案
针对大型数据集,优先采用矢量化操作和正则工具实现高效搜索,以下是满足需求的具体方案:
实现步骤
1. 示例数据准备
先构造模拟数据还原场景:
import pandas as pd import re data = { "Record ID": [1, 2, 3], "Text": [ "hello world hello python", "python is great, python is fun", "data science with python" ] } df = pd.DataFrame(data)
2. 核心搜索逻辑
定义函数提取目标词的词位置,再通过矢量化方法批量处理:
def get_target_positions(text, target): # 按空白分割文本为词列表(可根据需求调整分词规则,比如过滤标点) words = re.split(r'\s+', text.strip()) # 收集所有匹配目标词的位置(从1开始计数) return [idx + 1 for idx, word in enumerate(words) if word.lower() == target.lower()] # 指定目标词,批量处理Text列 target_word = "python" df["Word Positions"] = df["Text"].apply(lambda x: get_target_positions(x, target_word)) # 展开多值结果,保留Record ID result_df = df.explode("Word Positions").dropna(subset=["Word Positions"])
3. 大数据场景优化
如果处理百万级以上数据,可通过以下方式进一步提升效率:
- 使用
swifter库自动适配并行/矢量化执行,避免手动优化:
import swifter df["Word Positions"] = df["Text"].swifter.apply(lambda x: get_target_positions(x, target_word))
- 若需更精准的分词(如处理带标点的词汇),可提前加载
nltk分词器,避免重复初始化开销。
输出结果
最终result_df的结构如下:
| Record ID | Text | Word Positions |
|---|---|---|
| 1 | hello world hello python | 4 |
| 2 | python is great, python is fun | 1 |
| 2 | python is great, python is fun | 4 |
| 3 | data science with python | 4 |
关键注意事项
- 词位置计算以空白分割后的词索引+1为准,若需自定义分词规则(如去除标点),可修改
re.split的正则表达式。 - 代码默认大小写不敏感匹配,若需精确匹配,移除
lower()转换即可。 - 矢量化操作的效率远高于逐行循环,适合处理从数据库导入的大型DataFrame。
内容的提问来源于stack exchange,提问作者tshobe
相关产品推荐
相关产品推荐

