You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas如何筛选含唯一bigram且分数最高的关键词数据行

Pandas 实现基于二元词对的高分记录去重

核心逻辑

优先保留分数更高的关键词记录,维护一个全局已占用的二元词对集合:按分数从高到低遍历所有记录,如果当前记录的二元词对和已占用集合没有任何重叠,就保留这条记录,同时把它的所有二元词对加入已占用集合;如果存在重叠,直接过滤掉当前记录。

注:根据示例输出的匹配结果,这里的二元词对指单条关键词内任意两个词构成的无序对(和常规NLP中相邻二元组bigram的定义有区别,如果需要用相邻二元组,替换对应词对生成函数即可)

实现代码

import pandas as pd
from itertools import combinations

# 构造示例数据
df = pd.DataFrame({
    'key': [0,0,0,0,0],
    'keyword': [
        'ipsum dolor sit',
        'consectetur adipiscing elit',
        'ipsum dolor consectetur',
        'tellus id interdum',
        'ipsum interdum dolor'
    ],
    'score': [50,60,80,30,40]
})

# 生成单条文本的所有无序二元词对集合
def get_word_pairs(text):
    words = text.strip().split()
    # 对词排序保证无序对生成一致,去重避免重复词生成无效pair
    unique_words = sorted(list(set(words)))
    return set(combinations(unique_words, 2))

# 如果需要用传统相邻bigram,替换成下面这个函数即可
# def get_bigrams(text):
#     words = text.strip().split()
#     return set(zip(words, words[1:]))

# 按分数降序排序,高分记录优先获得保留资格
df_sorted = df.sort_values(by='score', ascending=False).reset_index(drop=False)
used_pairs = set()
keep_origin_idx = []

for _, row in df_sorted.iterrows():
    current_pairs = get_word_pairs(row['keyword'])
    # 无重叠词对则保留
    if not current_pairs & used_pairs:
        keep_origin_idx.append(row['index'])
        used_pairs.update(current_pairs)

# 筛选保留记录,按原索引排序还原原始顺序
result = df[df.index.isin(keep_origin_idx)].sort_index()
print(result)

运行结果

key                     keyword  score
1    0  consectetur adipiscing elit     60
2    0      ipsum dolor consectetur     80
3    0           tellus id interdum     30

扩展说明

如果你的数据需要按key字段分组独立去重,只要把上述遍历逻辑按key分组执行即可,每个分组维护独立的已占用词对集合。
如果需要使用相邻二元组(标准bigram)做去重判定,直接替换代码中注释的get_bigrams函数即可。

内容的提问来源于stack exchange,提问作者Sayandip Ghatak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 11:30:43