Pandas中count()、size()、nunique()的区别及分组调用差异
Pandas: Differences Between
count(), size(), and unique()/nunique() Methods Great question—these methods are easy to mix up, but their core differences lie in how they handle null values and whether they count total elements vs distinct values. Let’s break it down step by step.
1. Basic Differences on Series/DataFrame
count()
- Counts only non-null values (ignores
NaN,NaT, and other missing values). - For a Series: Returns the number of non-missing elements.
- For a DataFrame: Returns a Series where each entry is the count of non-null values in the corresponding column.
Example:
import pandas as pd import numpy as np s = pd.Series([1, 2, np.nan, 3]) print(s.count()) # Output: 3 (skips the NaN)
size()
- Returns the total number of elements, including null values.
- For a Series: Equivalent to
len(s)—gives the full length of the series regardless of missing data. - For a DataFrame: Returns the total number of cells (rows × columns), including all nulls.
Example:
s = pd.Series([1, 2, np.nan, 3]) print(s.size) # Output: 4 (includes the NaN)
unique()
- Returns an array of distinct values present in the Series (order matches their first occurrence).
- Includes
NaNvalues if they exist in the data. - Note: This method only works on Series (use
df.nunique()for DataFrames to get distinct counts per column).
Example:
s = pd.Series([1, 2, 2, np.nan, 3, np.nan]) print(s.unique()) # Output: array([ 1., 2., nan, 3.])
2. Differences When Used with groupby
When called on a GroupBy object, these methods operate per group. Your original example showed identical outputs because there were no nulls or duplicate values in data1—let’s modify the code to highlight the differences:
Modified Example Code:
import pandas as pd import numpy as np df = pd.DataFrame({ 'key1': ['a', 'a', 'b', 'b', 'a', 'a'], 'key2': ['one', 'two', 'one', 'two', 'one', 'three'], 'data1': [1.0, 1.0, np.nan, 3.0, np.nan, 2.0], 'data2': np.random.randn(6) }) grouped = df['data1'].groupby(df['key1']) print("Size per group:\n", grouped.size()) print("\nCount per group (non-null):\n", grouped.count()) print("\nNumber of unique values per group:\n", grouped.nunique())
Output:
Size per group: key1 a 4 b 2 Name: data1, dtype: int64 Count per group (non-null): key1 a 3 b 1 Name: data1, dtype: int64 Number of unique values per group: key1 a 2 b 1 Name: data1, dtype: int64
grouped.size()
- Returns the total number of rows in each group, including null values in the grouped series.
- In the example, group 'a' has 4 rows total (even with two
NaNs), group 'b' has 2 rows.
grouped.count()
- Returns the number of non-null values in each group for the target series.
- Group 'a' has 3 non-null values (two 1.0s and one 2.0), group 'b' has 1 non-null (3.0).
grouped.nunique()
- Returns the number of distinct non-null values per group (default
dropna=True—set toFalseto includeNaNas a unique value). - Group 'a' has 2 distinct non-null values, group 'b' has 1.
内容的提问来源于stack exchange,提问作者Eric Zhang
相关产品推荐
相关产品推荐

