You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从Pandas DataFrame文本列提取目标词左右各3个词?

解决方案

问题根源

你写的正则表达式直接把df['word']作为字符串字面量塞进了正则里,而不是动态替换成每行对应的目标词,这是导致全NaN的核心原因。另外原正则也没处理目标词多次出现、位于文本首尾或不存在的情况。

解决思路

用apply逐行处理文本,因为每行的目标词不同,需要针对每行单独计算保留范围:

  • 把文本拆成单词列表,定位所有目标词的位置
  • 计算需要保留的单词区间:从最左目标词的前3个词(不能小于0),到最右目标词的后3个词(不能超过单词列表长度)
  • 目标词不存在时返回空字符串

代码实现

import pandas as pd

# 构造示例DataFrame
data = {
    'id': [1,2,3,4],
    'text': [
        'i am working with john he is my colleague',
        'i watched the bond movie and the bond actor was amazing',
        'mary is my friend and we work together',
        'hello world hello python'
    ],
    'word': ['john', 'bond', 'mary', 'peter']
}
df = pd.DataFrame(data)

def get_retain_text(text, target_word):
    words = text.split()
    # 找到所有目标词的索引
    target_indices = [i for i, word in enumerate(words) if word == target_word]
    if not target_indices:
        return ''
    # 计算保留的起始和结束索引
    start_idx = max(0, min(target_indices) - 3)
    end_idx = min(len(words), max(target_indices) + 3 + 1)  # +1因为切片是左闭右开
    return ' '.join(words[start_idx:end_idx])

# 新增retain列
df['retain'] = df.apply(lambda row: get_retain_text(row['text'], row['word']), axis=1)

print(df)

输出结果

id                                                text    word                                              retain
0   1                i am working with john he is my colleague    john                  am working with john he is my
1   2  i watched the bond movie and the bond actor was amazing    bond  i watched the bond movie and the bond actor was amazing
2   3                   mary is my friend and we work together    mary                              mary is my friend
3   4                             hello world hello python   peter                                                     

内容的提问来源于stack exchange,提问作者MG Fern

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 12:52:16