如何在Python中可扩展地计算DataFrame二级子集的平均值?
Got it, let's break down how to solve this problem properly—no more manually typing family names for every subset you need to analyze.
First, let's start with a sample DataFrame that matches your scenario (easy to adapt to your actual data):
import pandas as pd # Sample data matching your use case data = { 'Family': ['Smith', 'Smith', 'Johnson', 'Johnson', 'Smith', 'Williams'], 'Gender': ['Male', 'Female', 'Female', 'Male', 'Female', 'Female'], 'Score': [80, 65, 75, 90, 100, 85] } df = pd.DataFrame(data)
1. Calculate Average for a Single Nested Subset (Scalable Function)
If you need to look up specific subsets (like Smith family females) but don't want to repeat code, wrap the logic in a reusable function. This way, you just pass the family and gender as parameters instead of hardcoding values every time:
def calculate_subset_average(df, family_name, gender): # Filter for the target family and gender, then compute mean score subset = df[(df['Family'] == family_name) & (df['Gender'] == gender)] return subset['Score'].mean() # Example: Get average score for Smith family females smith_female_avg = calculate_subset_average(df, 'Smith', 'Female') print(smith_female_avg) # Output: 82.5 (which is (65+100)/2)
2. Batch Calculate Averages for All Family-Gender Combinations (Best for Large Datasets)
For scenarios with lots of families, manually querying each one isn't feasible. Use groupby to automatically compute averages for every unique combination of family and gender—this is the most scalable approach:
# Group by Family and Gender, then calculate mean Score for each group all_subset_averages = df.groupby(['Family', 'Gender'])['Score'].mean().reset_index() print(all_subset_averages)
This will output a clean DataFrame with all your nested averages:
Family Gender Score 0 Johnson Female 75.0 1 Johnson Male 90.0 2 Smith Female 82.5 3 Smith Male 80.0 4 Williams Female 85.0
If you need to pull a specific value from this result later, you can use query to filter quickly:
# Extract average for Smith family females from the batch results smith_female_avg = all_subset_averages.query("Family == 'Smith' and Gender == 'Female'")['Score'].iloc[0]
3. Filter for Specific Families First (If Needed)
If you only care about a subset of families (but still too many to type manually), filter your DataFrame first using isin() before grouping:
# List of target families (can be as long as needed) target_families = ['Smith', 'Johnson'] # Filter the DataFrame to only include these families filtered_df = df[df['Family'].isin(target_families)] # Calculate averages for only the target families target_averages = filtered_df.groupby(['Family', 'Gender'])['Score'].mean().reset_index()
This approach keeps your code clean and adaptable, no matter how many families you need to analyze.
内容的提问来源于stack exchange,提问作者user6032266

