You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何查看DataFrame单一分类下所有值并统计标签数量

解决方法

问题分析

你当前的代码无法得到预期结果,核心问题有两个:

  1. 原始Tags列是逗号分隔的多标签字符串,并非单个标签,直接分组会把整个字符串当作一个标签统计
  2. 你需要的是每个分类下的总标签数,而非每组标签的出现次数

实现步骤与代码

1. 数据预处理:拆分多标签

先把Tags列里的逗号分隔标签拆成单独行,这是后续统计的基础。

2. 统计分类总标签数

计算每个Category下的标签总数。

3. 格式化输出样式

处理重复的分类名和总数,只在每组第一行显示。

完整代码:

import pandas as pd

# 构造原始DataFrame
data = {
    'Category': ['National', 'sports', 'Economy', 'sports'],
    'Tags': ['Politics, President', 'Football, tennis', 'Inflation', 'Cricket']
}
df = pd.DataFrame(data)

# 拆分多标签为单独行
df['Tags'] = df['Tags'].str.split(', ')
df_exploded = df.explode('Tags').reset_index(drop=True)

# 统计每个分类的标签总数
category_total = df_exploded.groupby('Category').size().reset_index(name='Count')

# 合并数据并格式化显示
result = df_exploded.merge(category_total, on='Category')
# 重复的分类名和总数置空
result['Category'] = result['Category'].mask(result['Category'] == result['Category'].shift(), '')
result['Count'] = result['Count'].mask(result['Count'] == result['Count'].shift(), '')

# 调整列顺序并打印
result = result[['Category', 'Tags', 'Count']]
print(result.to_string(index=False))

输出结果

Category            Tags           Count
sports              Football       3
                    tennis          
                    Cricket        

National            Politics       2
                    President       

Economy             Inflation      1

内容的提问来源于stack exchange,提问作者JATIN

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:55:23