求纯Pandas方案:保留多标签主题DataFrame的唯一完整路径主题
问题:纯Pandas实现层级标签DataFrame的去重保留逻辑
场景与需求
我们有一个多标签主题标注模型,会输出包含多层级标签的DataFrame。以3层级分类体系为例:
- 当某主题标注到L3时,会生成3行数据:一行包含L1、L2、L3,一行包含L1、L2、None,一行包含L1、None、None
- 若仅标注到L1或L2,则生成对应层级及更低层级补None的行
比如一个标注了l3_a(隶属于l1_a、l2_a)、l1_b、l2_b(隶属于l1_c)的条目,生成的DataFrame如下:
l1 l2 l3 1 l1_a l2_a l3_a 2 l1_a l2_a None 3 l1_a None None 4 l1_b None None 5 l1_c l2_b None 6 l1_c None None
我们需要保留最细粒度的唯一主题行:
- 保留第1行(L3级标注,可覆盖第2、3行)
- 保留第4行(仅L1级标注,无更细粒度行)
- 保留第5行(L2级标注,可覆盖第6行)
- 剔除第2、3、6行(被更细粒度的行覆盖)
当前循环实现方案
目前的解决方案通过循环从最高层级到最低层级逐步筛选,但不够优雅:
for i in range(n_levels, 0, -1): with_level = wide_topics[(wide_topics.iloc[:, :i].notnull()).all(axis=1)] if to_keep.empty: to_keep = with_level else: # Drop duplicates considering just these levels to_keep_this_level = ( pd.concat([to_keep.reset_index(), with_level.reset_index()]) .drop_duplicates(subset=with_level.columns[:i], keep=False) .set_index("index") ) # Finally, add the indexes to to_keep and drop duplicates (added 2 times from with_level) to_keep = ( pd.concat([to_keep_this_level, to_keep]) .reset_index() .drop_duplicates(keep="first") .set_index("index") )
寻求纯Pandas实现思路
上述方案可行,但依赖循环,希望找到无循环的纯Pandas实现方式,请问有什么思路?
内容的提问来源于stack exchange,提问作者Giovani Merlin
相关产品推荐
相关产品推荐

