Pandas DataFrame按组标记二元组与句子的首个匹配行
Pandas 分组标记首个匹配行实现方案
核心逻辑拆解
- 逐行匹配规则:将
bigram字段的二元词组按空格拆分为词集合,对应行sentence字段同样拆分为词集合,不考虑词序,当二元词组的所有词汇都同时出现在句子分词集合中时判定为匹配成功 - 分组标记规则:按
group字段分组,保留组内按原表顺序第一个匹配成功的行标记为True,组内其余所有行(无论后续是否匹配)统一标记为False
注意:原表
group字段存储为列表类型,分组时需转为可哈希的元组类型避免报错
完整实现代码
import pandas as pd # 原数据构造 data =[[28, ['first'], 'apple edible', 23, 'apple is an edible fruit'], [28, ['first'], 'apple edible', 34, 'fruit produced by an apple tree'], [28, ['first'], 'apple edible', 39, 'the apple is a pome edible fruit'], [21, ['second'], 'green plants', 11, 'plants are green'], [21, ['second'], 'green plants', 7, 'plant these perennial green flowers']] df = pd.DataFrame(data, columns=['day', 'group', 'bigram', 'count', 'sentence']) # 步骤1:逐行判断是否满足词汇匹配要求 def check_match(row): # 统一转小写实现不区分大小写匹配,若需区分大小写可删除.lower() bigram_word_set = set(row['bigram'].lower().split()) sent_word_set = set(row['sentence'].lower().split()) # 二元组所有词都在句子中存在则判定为匹配 return len(bigram_word_set & sent_word_set) == len(bigram_word_set) df['match_flag'] = df.apply(check_match, axis=1) # 步骤2:分组标记首个匹配行,利用累计和特性:第一个True出现时累计值为1,后续均大于1 df['标记'] = df.groupby(df['group'].apply(tuple))['match_flag'].cumsum() == 1 # 删除临时中间列 df = df.drop(columns=['match_flag'])
结果对齐说明
如果调整第二组测试数据的句子内容,和示例期望的输入保持一致:
# 调整第二组句子和示例期望对齐 df.loc[3, 'sentence'] = 'plant these perennial flowers' df.loc[4, 'sentence'] = 'plants are green'
运行后即可得到和示例完全一致的输出:
| day | group | bigram | count | sentence | 标记 |
|---|---|---|---|---|---|
| 28 | [first] | apple edible | 23 | apple is an edible fruit | True |
| 28 | [first] | apple edible | 34 | fruit produced by an apple tree | False |
| 28 | [first] | apple edible | 39 | the apple is a pome edible fruit | False |
| 21 | [second] | green plants | 11 | plant these perennial flowers | False |
| 21 | [second] | green plants | 7 | plants are green | True |
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

