如何从Pandas DataFrame文本列提取目标词左右各3个词?
解决方案
问题根源
你写的正则表达式直接把df['word']作为字符串字面量塞进了正则里,而不是动态替换成每行对应的目标词,这是导致全NaN的核心原因。另外原正则也没处理目标词多次出现、位于文本首尾或不存在的情况。
解决思路
用apply逐行处理文本,因为每行的目标词不同,需要针对每行单独计算保留范围:
- 把文本拆成单词列表,定位所有目标词的位置
- 计算需要保留的单词区间:从最左目标词的前3个词(不能小于0),到最右目标词的后3个词(不能超过单词列表长度)
- 目标词不存在时返回空字符串
代码实现
import pandas as pd # 构造示例DataFrame data = { 'id': [1,2,3,4], 'text': [ 'i am working with john he is my colleague', 'i watched the bond movie and the bond actor was amazing', 'mary is my friend and we work together', 'hello world hello python' ], 'word': ['john', 'bond', 'mary', 'peter'] } df = pd.DataFrame(data) def get_retain_text(text, target_word): words = text.split() # 找到所有目标词的索引 target_indices = [i for i, word in enumerate(words) if word == target_word] if not target_indices: return '' # 计算保留的起始和结束索引 start_idx = max(0, min(target_indices) - 3) end_idx = min(len(words), max(target_indices) + 3 + 1) # +1因为切片是左闭右开 return ' '.join(words[start_idx:end_idx]) # 新增retain列 df['retain'] = df.apply(lambda row: get_retain_text(row['text'], row['word']), axis=1) print(df)
输出结果
id text word retain 0 1 i am working with john he is my colleague john am working with john he is my 1 2 i watched the bond movie and the bond actor was amazing bond i watched the bond movie and the bond actor was amazing 2 3 mary is my friend and we work together mary mary is my friend 3 4 hello world hello python peter
内容的提问来源于stack exchange,提问作者MG Fern
相关产品推荐
相关产品推荐

