You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组子集二次分组返回已过滤值的问题及解决

问题原因及解决方法

原因

Pandas的pd.Categorical类型会保留所有原始定义的类别,哪怕你过滤掉了对应类别的行,后续groupby操作默认会包含这些没有数据的空类别。而Plotly Express的px.bar内部调用groupby.get_group时,会遵循这个规则,所以即使数据里没有C组,依然会显示这个空组。

解决方法

1. 移除未使用的类别

过滤数据后,对Categorical列执行remove_unused_categories(),只保留当前数据中存在的类别:

# 假设原始数据已过滤掉C组
df['category'] = df['category'].cat.remove_unused_categories()
# 再调用px.bar就不会出现C组了
fig = px.bar(df, x='category', y='value')

2. Groupby时指定observed=True

在groupby操作中添加observed=True参数(Pandas 1.1.0及以上版本支持),强制只保留有实际数据的类别:

grouped_df = df.groupby('category', observed=True).agg({'value': 'sum'})
fig = px.bar(grouped_df, x=grouped_df.index, y='value')

3. 手动限定Plotly的显示类别

在px.bar中通过category_orders参数直接指定要显示的类别列表,忽略空类别:

fig = px.bar(df, x='category', y='value', category_orders={'category': ['A', 'B']})

内容的提问来源于stack exchange,提问作者Marco Aurélio Guerra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 02:35:23