You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas内置函数基于词汇匹配数匹配主题模型的主题

用Pandas匹配跨数据集的主题词汇

你可以用Pandas的向量化操作替代手动遍历,核心是通过集合交集计算词汇匹配数量,再找到每个topics1主题对应的最优匹配主题——不需要依赖eq(),它只适合元素级相等判断,不适用集合场景。

实现代码

import pandas as pd

# 示例数据
topics1 = pd.DataFrame({'topic_num':[1,2,3],
                        'words':[['red','blue','green'],
                                 ['blue','sky','cloud'],
                                 ['eat','food','nomnom']]
                        })
topics2 = pd.DataFrame({'topic_num':[1,2,3],
                        'words':[['blue','sky','airplane'],
                                 ['blue','green','yellow'],
                                 ['mac','bit','byte']]
                        })

# 将topics2的词汇转为集合,提升交集计算效率
topics2['word_set'] = topics2['words'].apply(set)

# 定义函数:计算当前主题与topics2所有主题的词汇匹配数,返回最优匹配的主题编号
def get_best_match(word_list):
    current_set = set(word_list)
    # 计算与每个topics2主题的交集大小
    match_counts = topics2['word_set'].apply(lambda x: len(current_set & x))
    # 返回匹配数最多的topics2主题编号(若有多个最大值,取第一个)
    return topics2.loc[match_counts.idxmax(), 'topic_num']

# 为topics1每个主题匹配最优主题
topics1['matched_topic_num'] = topics1['words'].apply(get_best_match)

print(topics1)

输出结果

topic_num               words  matched_topic_num
0          1  [red, blue, green]                  2
1          2  [blue, sky, cloud]                  1
2          3   [eat, food, nomnom]                  1

说明

  • 用集合操作计算交集比逐词遍历更高效,尤其是主题词汇量较大时。
  • 若存在多个topics2主题与目标主题匹配数相同,idxmax()会返回第一个出现的主题,如需处理这种平局情况,可以添加额外逻辑(比如返回所有匹配主题编号)。

内容的提问来源于stack exchange,提问作者Adam_G

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 13:53:25