You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多分组变量下统计多个变量的缺失值数量

Fixing Multi-Group Missing Value Counting in Pandas

Hey there! Let's break down why you're hitting that ValueError and get your multi-group missing value stats working smoothly.

Why Your Original Code Failed

The error Grouper for '<class 'pandas.core.frame.DataFrame'>' not 1-dimensional pops up because you passed an entire multi-column DataFrame (df[['type', 'type2']]) as the grouping key to groupby(). Pandas expects grouping keys to be 1-dimensional—like a single Series, list of column names, or list of individual Series—not a full DataFrame. That's where your code got stuck.

Simple, Working Solutions

Here are two clean ways to calculate missing values across multiple groups:

Method 1: Group by Column Names Directly (Concise)

Use the original DataFrame's groupby() method with your list of grouping columns, then target the variables you want to check and compute missing values in one chain:

import pandas as pd

# Your sample data
df = pd.DataFrame({
    'type': ['A', 'A', 'B'], 
    'type2': ['Z', 'Y', 'Y'], 
    'var1': [1, None, 2], 
    'var2': [None, None, 3]
})

grouping_vars = ['type', 'type2']
vars_to_check = ['var1', 'var2']

# Calculate missing values per group
result = df.groupby(grouping_vars, sort=False)[vars_to_check].apply(lambda x: x.isna().sum()).reset_index()
print(result)

Method 2: Explicit Missing Value Flags (Intuitive)

If you prefer clarity over brevity, create new columns that flag missing values, then group and sum those flags:

# Add columns to mark missing values
df['var1_missing'] = df['var1'].isna()
df['var2_missing'] = df['var2'].isna()

# Sum the flags for each group
result = df.groupby(grouping_vars, sort=False)[['var1_missing', 'var2_missing']].sum().reset_index()
print(result)

Expected Output

Both methods will return this correct result:

type type2  var1  var2
0    A     Z     0     1
1    A     Y     1     1
2    B     Y     0     0

Key Takeaways

  • Stick to passing lists of column names to groupby() for multi-group tasks—it's the standard pandas pattern.
  • apply(lambda x: x.isna().sum()) lets you compute missing values directly within the groupby workflow.
  • Explicit missing value flags are great if you need to reuse those indicators for other analyses later.

内容的提问来源于stack exchange,提问作者Helen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 13:27:45