You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中count()、size()、nunique()的区别及分组调用差异

Pandas: Differences Between count(), size(), and unique()/nunique() Methods

Great question—these methods are easy to mix up, but their core differences lie in how they handle null values and whether they count total elements vs distinct values. Let’s break it down step by step.


1. Basic Differences on Series/DataFrame

count()

  • Counts only non-null values (ignores NaN, NaT, and other missing values).
  • For a Series: Returns the number of non-missing elements.
  • For a DataFrame: Returns a Series where each entry is the count of non-null values in the corresponding column.

Example:

import pandas as pd
import numpy as np

s = pd.Series([1, 2, np.nan, 3])
print(s.count())  # Output: 3 (skips the NaN)

size()

  • Returns the total number of elements, including null values.
  • For a Series: Equivalent to len(s)—gives the full length of the series regardless of missing data.
  • For a DataFrame: Returns the total number of cells (rows × columns), including all nulls.

Example:

s = pd.Series([1, 2, np.nan, 3])
print(s.size)  # Output: 4 (includes the NaN)

unique()

  • Returns an array of distinct values present in the Series (order matches their first occurrence).
  • Includes NaN values if they exist in the data.
  • Note: This method only works on Series (use df.nunique() for DataFrames to get distinct counts per column).

Example:

s = pd.Series([1, 2, 2, np.nan, 3, np.nan])
print(s.unique())  # Output: array([ 1.,  2., nan,  3.])

2. Differences When Used with groupby

When called on a GroupBy object, these methods operate per group. Your original example showed identical outputs because there were no nulls or duplicate values in data1—let’s modify the code to highlight the differences:

Modified Example Code:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'key1': ['a', 'a', 'b', 'b', 'a', 'a'],
    'key2': ['one', 'two', 'one', 'two', 'one', 'three'],
    'data1': [1.0, 1.0, np.nan, 3.0, np.nan, 2.0],
    'data2': np.random.randn(6)
})

grouped = df['data1'].groupby(df['key1'])
print("Size per group:\n", grouped.size())
print("\nCount per group (non-null):\n", grouped.count())
print("\nNumber of unique values per group:\n", grouped.nunique())

Output:

Size per group:
 key1
a    4
b    2
Name: data1, dtype: int64

Count per group (non-null):
 key1
a    3
b    1
Name: data1, dtype: int64

Number of unique values per group:
 key1
a    2
b    1
Name: data1, dtype: int64

grouped.size()

  • Returns the total number of rows in each group, including null values in the grouped series.
  • In the example, group 'a' has 4 rows total (even with two NaNs), group 'b' has 2 rows.

grouped.count()

  • Returns the number of non-null values in each group for the target series.
  • Group 'a' has 3 non-null values (two 1.0s and one 2.0), group 'b' has 1 non-null (3.0).

grouped.nunique()

  • Returns the number of distinct non-null values per group (default dropna=True—set to False to include NaN as a unique value).
  • Group 'a' has 2 distinct non-null values, group 'b' has 1.

内容的提问来源于stack exchange,提问作者Eric Zhang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:12:24