You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

不同标签集多标签分类模型预测共现计数矩阵Python实现问询

多标签分类结果共现计数矩阵实现方案

实现思路

  • 先明确两个模型各自的全量标签集合,避免标签混淆
  • 逐张拆分每张图像归属两个模型的预测标签
  • 对单张图所有模型1标签和模型2标签的两两组合进行计数
  • 最终输出行对应模型1标签、列对应模型2标签的计数矩阵

完整可运行代码

import pandas as pd

# 1. 替换为你实际的两个模型全量标签列表
model1_labels = ["m1", "m2"]  # 第一个模型共28个标签,此处为示例值
model2_labels = ["t1", "t2", "t3"]  # 第二个模型共120个标签,此处为示例值

# 2. 替换为你所有图像的预测结果,每个元素对应单张图的所有预测标签
all_img_preds = [
    ['m1', 'm2', 't1'],
    ['m1', 't1'],
    ['m1', 'm2', 't1', 't3']
]

# 3. 初始化全0计数矩阵,自动对齐所有标签,缺失共现默认填0
count_matrix = pd.DataFrame(0, index=model1_labels, columns=model2_labels)

# 4. 逐样本统计共现次数
for img_tags in all_img_preds:
    # 拆分当前图归属两个模型的标签
    current_m_tags = [tag for tag in img_tags if tag in model1_labels]
    current_t_tags = [tag for tag in img_tags if tag in model2_labels]
    # 两两组合计数+1
    for m_tag in current_m_tags:
        for t_tag in current_t_tags:
            count_matrix.loc[m_tag, t_tag] += 1

# 打印输出结果
print(count_matrix)

运行输出结果和你要求的格式完全一致:

t1  t2  t3
m1   3   0   1
m2   2   0   1

性能说明

针对你提到的1200张图像、28个模型1标签、120个模型2标签的业务场景,该方案时间复杂度为O(样本数 * 单样本模型1标签数 * 单样本模型2标签数),实际运行耗时不足1秒,完全满足处理需求。

如果偏好使用pandas.crosstab实现,可使用如下替代代码,效果完全一致:

pairs = []
for img_tags in all_img_preds:
    ms = [t for t in img_tags if t in model1_labels]
    ts = [t for t in img_tags if t in model2_labels]
    pairs.extend([(m, t) for m in ms for t in ts])

count_matrix = pd.crosstab(
    [p[0] for p in pairs],
    [p[1] for p in pairs]
).reindex(
    index=model1_labels,
    columns=model2_labels,
    fill_value=0
)

内容的提问来源于stack exchange,提问作者JBVasc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:36:03