如何在Pandas中计算GroupBy分组非空值计数的中位数?
Calculate Median of Non-Empty Value Counts per Group in Pandas
Here's a straightforward way to solve this problem in Pandas, broken down into simple steps:
Step 1: Count Non-Empty Values per Group
First, group your DataFrame by the g column. For each group, count how many non-empty entries exist in the val column—we can do this by checking which values aren't empty strings (x != '') and summing those boolean results (since True maps to 1 and False maps to 0).
Step 2: Compute the Median of These Counts
Once you have the non-empty count for each group, call the .median() method on the resulting Series to get the median value across all groups.
Full Code Example
Let’s use your sample DataFrame to demonstrate:
import pandas as pd # Create your sample DataFrame df = pd.DataFrame({ 'g': [1, 1, 2, 2, 2, 3], 'val': ['a', '', 'b', '', 'c', ''] }) # Step 1: Calculate non-empty value counts for each group non_empty_counts = df.groupby('g')['val'].apply(lambda x: (x != '').sum()) # Step 2: Find the median of these counts median_result = non_empty_counts.median() # Print results to verify print("Non-empty counts per group:") print(non_empty_counts) # Output: # g # 1 1 # 2 2 # 3 0 # Name: val, dtype: int64 print("\nMedian of non-empty counts:", median_result) # Output: Median of non-empty counts: 1.0
Alternative Concise Syntax
You can also use .agg() instead of .apply() for a more compact version:
median_result = df.groupby('g')['val'].agg(lambda x: (x != '').sum()).median()
Key Notes
- This approach automatically includes groups with zero non-empty values (like group 3 in your example), so your median calculation accounts for all groups in the original DataFrame.
- If your "empty" values are
NaNinstead of empty strings, replacex != ''withx.notna().sum()to count non-null values instead.
内容的提问来源于stack exchange,提问作者mkmostafa
相关产品推荐
相关产品推荐

