You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Pandas DataFrame中提取数组元素间的介词与连词并生成新列

如何从Pandas DataFrame的token序列中提取特定连接词汇?

我有一个如下的Pandas DataFrame,其中token_1列的列表不包含介词或连词,我想从token_2列的列表中提取token_1元素匹配点之间、长度小于4的词汇,最终生成new_tokens列。

示例DataFrame代码:

import pandas as pd
data = {'token_1': [['cat', 'bag', 'sitting'], ['dog', 'eats', 'bowls'], ['mouse', 'mustache', 'tail'], ['dog', 'eat', 'meat']], 'token_2': [['cat', 'from', 'bag', 'cat', 'in', 'bag', 'sitting', 'whole', 'day'], ['dog', 'eats', 'from', 'bowls', 'dog', 'eats', 'always', 'from', 'bowls', 'eats', 'bowl'], ['mouse', 'with', 'a', 'big', 'tail', 'and','ears', 'a', 'mouse', 'with', 'a', 'mustache', 'and', 'a', 'tail' ,'runs', 'fast'], ['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat']]}
df = pd.DataFrame(data)

期望生成的new_tokens列结果如下:

token_1new_tokenstoken_2
['cat', 'bag', 'sitting']['cat', 'in', 'bag', 'sitting']['cat', 'from', 'bag', 'cat', 'in', 'bag', 'sitting', 'whole', 'day']
['dog', 'eats', 'bowls']['dog', 'eats', 'from', 'bowls']['dog', 'eats', 'from', 'bowls', 'dog', 'eats', 'always', 'from', 'bowls', 'eats', 'bowl']
['mouse', 'mustache', 'tail']['mouse', 'with', 'mustache', 'and', 'tail']['mouse', 'with', 'a', 'big', 'tail', 'and','ears', 'a', 'mouse', 'with', 'a', 'mustache', 'and', 'a', 'tail' ,'runs', 'fast']
['dog', 'eat', 'meat']['dog', 'eat', 'meat']['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat']

我初步梳理了一些步骤,但想知道有没有更简便的实现方法?


简便实现方法

你可以通过自定义函数结合pandas.DataFrame.apply来高效实现这个需求,核心思路是定位token_1中相邻元素在token_2中的连续匹配区间,提取中间符合长度要求的词汇,具体代码如下:

def process_tokens(row):
    tokens1 = row['token_1']
    tokens2 = row['token_2']
    new_tokens = []
    prev_pos = -1
    
    # 遍历token_1中的每个元素对(当前元素+下一个元素)
    for i in range(len(tokens1)):
        current_token = tokens1[i]
        # 找到当前token在token_2中所有出现的位置
        current_positions = [idx for idx, t in enumerate(tokens2) if t == current_token]
        
        # 处理第一个元素:取第一个出现的位置,加入结果
        if i == 0:
            if current_positions:
                prev_pos = current_positions[0]
                new_tokens.append(current_token)
            continue
        
        # 找到当前token在prev_pos之后的第一个出现位置
        next_pos = None
        for pos in current_positions:
            if pos > prev_pos:
                next_pos = pos
                break
        
        if next_pos is None:
            # 找不到后续匹配,直接加入当前token
            new_tokens.append(current_token)
            prev_pos = len(tokens2) - 1
            continue
        
        # 提取prev_pos+1到next_pos-1之间的词汇,筛选长度<4的
        between_tokens = tokens2[prev_pos+1 : next_pos]
        filtered = [t for t in between_tokens if len(t) < 4]
        
        # 将筛选后的词汇和当前token加入结果
        new_tokens.extend(filtered)
        new_tokens.append(current_token)
        
        # 更新prev_pos为当前token的位置
        prev_pos = next_pos
    
    return new_tokens

# 生成new_tokens列
df['new_tokens'] = df.apply(process_tokens, axis=1)

代码说明

  1. 定位匹配位置:对于token_1中的每个元素,先找到它在token_2中的所有出现位置;
  2. 区间提取:针对相邻的两个token_1元素,找到它们在token_2中连续的匹配区间(前一个元素的位置之后的第一个当前元素位置);
  3. 筛选词汇:提取区间内长度小于4的词汇,加入结果列表;
  4. 边界处理:如果找不到后续匹配的元素,直接将当前token加入结果,避免报错。

运行这段代码后,就能得到你期望的new_tokens列结果。

内容的提问来源于stack exchange,提问作者Rory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 01:22:36