在Pandas DataFrame中提取数组元素间的介词与连词并生成新列
如何从Pandas DataFrame的token序列中提取特定连接词汇?
我有一个如下的Pandas DataFrame,其中token_1列的列表不包含介词或连词,我想从token_2列的列表中提取token_1元素匹配点之间、长度小于4的词汇,最终生成new_tokens列。
示例DataFrame代码:
import pandas as pd data = {'token_1': [['cat', 'bag', 'sitting'], ['dog', 'eats', 'bowls'], ['mouse', 'mustache', 'tail'], ['dog', 'eat', 'meat']], 'token_2': [['cat', 'from', 'bag', 'cat', 'in', 'bag', 'sitting', 'whole', 'day'], ['dog', 'eats', 'from', 'bowls', 'dog', 'eats', 'always', 'from', 'bowls', 'eats', 'bowl'], ['mouse', 'with', 'a', 'big', 'tail', 'and','ears', 'a', 'mouse', 'with', 'a', 'mustache', 'and', 'a', 'tail' ,'runs', 'fast'], ['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat']]} df = pd.DataFrame(data)
期望生成的new_tokens列结果如下:
| token_1 | new_tokens | token_2 |
|---|---|---|
| ['cat', 'bag', 'sitting'] | ['cat', 'in', 'bag', 'sitting'] | ['cat', 'from', 'bag', 'cat', 'in', 'bag', 'sitting', 'whole', 'day'] |
| ['dog', 'eats', 'bowls'] | ['dog', 'eats', 'from', 'bowls'] | ['dog', 'eats', 'from', 'bowls', 'dog', 'eats', 'always', 'from', 'bowls', 'eats', 'bowl'] |
| ['mouse', 'mustache', 'tail'] | ['mouse', 'with', 'mustache', 'and', 'tail'] | ['mouse', 'with', 'a', 'big', 'tail', 'and','ears', 'a', 'mouse', 'with', 'a', 'mustache', 'and', 'a', 'tail' ,'runs', 'fast'] |
| ['dog', 'eat', 'meat'] | ['dog', 'eat', 'meat'] | ['dog', 'eat', 'meat', 'chicken', 'from', 'bowl','dog','see','meat','eat'] |
我初步梳理了一些步骤,但想知道有没有更简便的实现方法?
简便实现方法
你可以通过自定义函数结合pandas.DataFrame.apply来高效实现这个需求,核心思路是定位token_1中相邻元素在token_2中的连续匹配区间,提取中间符合长度要求的词汇,具体代码如下:
def process_tokens(row): tokens1 = row['token_1'] tokens2 = row['token_2'] new_tokens = [] prev_pos = -1 # 遍历token_1中的每个元素对(当前元素+下一个元素) for i in range(len(tokens1)): current_token = tokens1[i] # 找到当前token在token_2中所有出现的位置 current_positions = [idx for idx, t in enumerate(tokens2) if t == current_token] # 处理第一个元素:取第一个出现的位置,加入结果 if i == 0: if current_positions: prev_pos = current_positions[0] new_tokens.append(current_token) continue # 找到当前token在prev_pos之后的第一个出现位置 next_pos = None for pos in current_positions: if pos > prev_pos: next_pos = pos break if next_pos is None: # 找不到后续匹配,直接加入当前token new_tokens.append(current_token) prev_pos = len(tokens2) - 1 continue # 提取prev_pos+1到next_pos-1之间的词汇,筛选长度<4的 between_tokens = tokens2[prev_pos+1 : next_pos] filtered = [t for t in between_tokens if len(t) < 4] # 将筛选后的词汇和当前token加入结果 new_tokens.extend(filtered) new_tokens.append(current_token) # 更新prev_pos为当前token的位置 prev_pos = next_pos return new_tokens # 生成new_tokens列 df['new_tokens'] = df.apply(process_tokens, axis=1)
代码说明
- 定位匹配位置:对于token_1中的每个元素,先找到它在token_2中的所有出现位置;
- 区间提取:针对相邻的两个token_1元素,找到它们在token_2中连续的匹配区间(前一个元素的位置之后的第一个当前元素位置);
- 筛选词汇:提取区间内长度小于4的词汇,加入结果列表;
- 边界处理:如果找不到后续匹配的元素,直接将当前token加入结果,避免报错。
运行这段代码后,就能得到你期望的new_tokens列结果。
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

