分组数据均值计算:基于员工薪资DataFrame的技术咨询
Got it, let's walk through how to calculate grouped means for your employee salary DataFrame. I’m assuming you’re using pandas since you’re working with a DataFrame—here are the most common scenarios you’ll probably need:
For example, if you want to see average phd years, service years, and salary broken down by employee rank:
import pandas as pd # Replace 'df' with your actual DataFrame name rank_grouped_means = df.groupby('rank').mean() print(rank_grouped_means)
This will automatically compute the mean for every numeric column (phd, srv, salary) grouped by each unique value in the rank column (like Prof, Assoc, Asst).
If you want to drill down further—say, average salary grouped by both rank and gender:
rank_sex_salary_mean = df.groupby(['rank', 'sex'])['salary'].mean().reset_index() print(rank_sex_salary_mean)
The reset_index() converts the grouped fields from index columns back to regular columns, making the output easier to read. We also specify ['salary'] to only calculate the mean for that column instead of all numeric fields.
If you only care about average salary and phd years grouped by department (dsc):
dept_specific_means = df.groupby('dsc')[['salary', 'phd']].mean() print(dept_specific_means)
By passing a list of column names (['salary', 'phd']), we limit the calculation to just those two fields.
If your dataset has missing entries, you can either clean the data first or explicitly tell pandas to skip NaNs (which it does by default, but it’s good to be explicit):
# Option 1: Drop rows with any missing values before grouping clean_df = df.dropna() clean_grouped_means = clean_df.groupby('rank')['salary'].mean() # Option 2: Explicitly skip NaNs during calculation (default behavior) grouped_means_skipna = df.groupby('rank')['salary'].mean(skipna=True)
To make the mean values easier to read, use round() to limit decimal places:
rank_salary_mean_rounded = df.groupby('rank')['salary'].mean().round(2) print(rank_salary_mean_rounded)
内容的提问来源于stack exchange,提问作者Dan C

