Pandas DataFrame按指定顺序匹配词元保留索引生成新列
实现方法
核心思路是逐行遍历数据,在每行的token_pos元组序列中,用固定长度滑动窗口匹配和token_1词组长度一致、词顺序完全对应的连续片段,取第一个匹配到的片段作为新列值。
完整代码
import pandas as pd # 原始数据构造 data = {'token_1': [['cat', 'run','today'],['dog', 'eat', 'meat']], 'token_2': [['cat', 'in', 'the' , 'morning','cat', 'run', 'today', 'very', 'quick', 'cat','today', 'jump', 'and', 'run', 'run', 'cat', 'today'],['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat']], 'token_pos':[[(0,'cat'), (2,'the') , (3,'morning'), (4,'cat'), (5,'run'), (6,'today'), (7,'very'), (8,'quick'), (9,'cat'), (10,'today'), (12,'and'), (13,'run'), (14,'run'), (15,'cat'), (16,'today')], [(0,'dog'), (1,'eat'), (2,'meat'), (3,'chicken'),(5,'bowl'),(6,'dog'),(7,'see'),(8,'meat'),(9,'eat'),(15,'bowl')]]} df = pd.DataFrame(data) # 定义逐行匹配函数 def match_continuous_seq(row): target_words = row['token_1'] window_size = len(target_words) pos_tuples = row['token_pos'] # 滑动窗口遍历所有可能的连续片段 for start_idx in range(len(pos_tuples) - window_size + 1): window = pos_tuples[start_idx : start_idx + window_size] # 提取窗口内的词和目标序列比对 window_words = [item[1] for item in window] if window_words == target_words: return window # 无匹配时返回空列表,可按需调整 return [] # 生成新列 df['new_token'] = df.apply(match_continuous_seq, axis=1) # 输出结果验证 print(df[['new_token']])
运行输出
new_token 0 [(4, 'cat'), (5, 'run'), (6, 'today')] 1 [(0, 'dog'), (1, 'eat'), (2, 'meat')]
注意事项
- 逻辑默认
token_pos内的元组已按位置索引升序排列,和输入数据结构一致 - 匹配时返回第一个符合顺序要求的连续词元片段,和需求预期对齐
- 若某行不存在匹配片段,函数默认返回空列表,可根据实际场景调整返回规则
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

