You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame按指定顺序匹配词元保留索引生成新列

实现方法

核心思路是逐行遍历数据,在每行的token_pos元组序列中,用固定长度滑动窗口匹配和token_1词组长度一致、词顺序完全对应的连续片段,取第一个匹配到的片段作为新列值。

完整代码

import pandas as pd

# 原始数据构造
data = {'token_1': [['cat', 'run','today'],['dog', 'eat', 'meat']],
        'token_2': [['cat', 'in', 'the' , 'morning','cat', 'run', 'today',
                      'very', 'quick', 'cat','today', 'jump', 'and', 'run', 'run', 'cat', 'today'],['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat']],
       'token_pos':[[(0,'cat'), (2,'the') , (3,'morning'), (4,'cat'), (5,'run'), (6,'today'),
                      (7,'very'), (8,'quick'), (9,'cat'), (10,'today'), (12,'and'), (13,'run'), (14,'run'), (15,'cat'), (16,'today')],
                    [(0,'dog'), (1,'eat'), (2,'meat'), (3,'chicken'),(5,'bowl'),(6,'dog'),(7,'see'),(8,'meat'),(9,'eat'),(15,'bowl')]]}
    
df = pd.DataFrame(data)

# 定义逐行匹配函数
def match_continuous_seq(row):
    target_words = row['token_1']
    window_size = len(target_words)
    pos_tuples = row['token_pos']
    # 滑动窗口遍历所有可能的连续片段
    for start_idx in range(len(pos_tuples) - window_size + 1):
        window = pos_tuples[start_idx : start_idx + window_size]
        # 提取窗口内的词和目标序列比对
        window_words = [item[1] for item in window]
        if window_words == target_words:
            return window
    # 无匹配时返回空列表,可按需调整
    return []

# 生成新列
df['new_token'] = df.apply(match_continuous_seq, axis=1)

# 输出结果验证
print(df[['new_token']])

运行输出

new_token
0  [(4, 'cat'), (5, 'run'), (6, 'today')]
1  [(0, 'dog'), (1, 'eat'), (2, 'meat')]

注意事项

  • 逻辑默认token_pos内的元组已按位置索引升序排列,和输入数据结构一致
  • 匹配时返回第一个符合顺序要求的连续词元片段,和需求预期对齐
  • 若某行不存在匹配片段,函数默认返回空列表,可根据实际场景调整返回规则

内容的提问来源于stack exchange,提问作者Rory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 21:54:32