You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas DataFrame按组标记二元组与句子的首个匹配行

Pandas 分组标记首个匹配行实现方案

核心逻辑拆解

  • 逐行匹配规则:将bigram字段的二元词组按空格拆分为词集合,对应行sentence字段同样拆分为词集合,不考虑词序,当二元词组的所有词汇都同时出现在句子分词集合中时判定为匹配成功
  • 分组标记规则:按group字段分组,保留组内按原表顺序第一个匹配成功的行标记为True,组内其余所有行(无论后续是否匹配)统一标记为False

注意:原表group字段存储为列表类型,分组时需转为可哈希的元组类型避免报错

完整实现代码

import pandas as pd

# 原数据构造
data =[[28, ['first'], 'apple edible', 23, 'apple is an edible fruit'],
 [28, ['first'], 'apple edible', 34, 'fruit produced by an apple tree'],
 [28, ['first'], 'apple edible', 39, 'the apple is a pome edible fruit'],
 [21, ['second'], 'green plants', 11, 'plants are green'],
 [21, ['second'], 'green plants', 7, 'plant these perennial green flowers']]
df = pd.DataFrame(data, columns=['day', 'group',  'bigram', 'count', 'sentence'])

# 步骤1:逐行判断是否满足词汇匹配要求
def check_match(row):
    # 统一转小写实现不区分大小写匹配,若需区分大小写可删除.lower()
    bigram_word_set = set(row['bigram'].lower().split())
    sent_word_set = set(row['sentence'].lower().split())
    # 二元组所有词都在句子中存在则判定为匹配
    return len(bigram_word_set & sent_word_set) == len(bigram_word_set)

df['match_flag'] = df.apply(check_match, axis=1)

# 步骤2:分组标记首个匹配行,利用累计和特性:第一个True出现时累计值为1,后续均大于1
df['标记'] = df.groupby(df['group'].apply(tuple))['match_flag'].cumsum() == 1

# 删除临时中间列
df = df.drop(columns=['match_flag'])

结果对齐说明

如果调整第二组测试数据的句子内容,和示例期望的输入保持一致:

# 调整第二组句子和示例期望对齐
df.loc[3, 'sentence'] = 'plant these perennial flowers'
df.loc[4, 'sentence'] = 'plants are green'

运行后即可得到和示例完全一致的输出:

daygroupbigramcountsentence标记
28[first]apple edible23apple is an edible fruitTrue
28[first]apple edible34fruit produced by an apple treeFalse
28[first]apple edible39the apple is a pome edible fruitFalse
21[second]green plants11plant these perennial flowersFalse
21[second]green plants7plants are greenTrue

内容的提问来源于stack exchange,提问作者Rory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 11:36:18