如何使用Python Pandas对数据集中的个体进行分组
Alright, let's walk through exactly how to group your dataset by the Subject column using Pandas. First, I noticed your Variable3 list was incomplete (only 5 elements instead of 12), so I'll fix that first to ensure we can work with a valid, runnable DataFrame.
Step 1: Fix and Load the Dataset
First, let's get our data into a proper Pandas DataFrame:
import pandas as pd # Corrected full dataset (filled in missing Variable3 values) d1 = { 'Subject': ['Subject1','Subject1','Subject1','Subject2','Subject2','Subject2','Subject3','Subject3','Subject3','Subject4','Subject4','Subject4'], 'Event':['1','2','3','1','2','3','1','2','3','1','2','3'], 'Category':['1','1','2','2','1','2','2','','2','1','1',''], 'Variable1':['1','2','3','4','5','6','7','8','9','10','11','12'], 'Variable2':['12','11','10','9','8','7','6','5','4','3','2','1'], 'Variable3': ['-6','-5','-4','-3','-2','-1','0','1','2','3','4','5'] } df = pd.DataFrame(d1)
Step 2: Basic Grouping by Subject
The core of grouping is Pandas' groupby() method. This creates a DataFrameGroupBy object that we can manipulate for all sorts of tasks:
# Group the DataFrame by the 'Subject' column subject_groups = df.groupby('Subject')
Step 3: Common Operations on Grouped Data
Now that we have our groups, here are the most useful things you can do with them:
1. Iterate through groups to inspect data
If you want to see exactly what's in each subject's group, loop through the DataFrameGroupBy object:
for subject_name, group_data in subject_groups: print(f"--- Data for {subject_name} ---") print(group_data) print("\n")
2. Aggregate statistics (sum, mean, max, etc.)
Use agg() to compute multiple stats at once for each group. For example, let's calculate the mean and sum of Variable1, plus the max of Variable2 per subject:
# Aggregate multiple metrics group_stats = subject_groups.agg( Variable1_average=('Variable1', 'mean'), Variable1_total=('Variable1', 'sum'), Variable2_highest=('Variable2', 'max') ) print(group_stats)
Note: Since your Variable1/Variable2 are stored as strings, Pandas will convert them to numeric automatically for these calculations, but it's safer to explicitly cast them with astype(int) if you're working with larger datasets.
3. Apply custom functions to groups
For more complex logic, use apply() to run your own function on each group. Let's calculate the total difference between Variable2 and Variable1 for each subject:
def calculate_total_difference(group): # Convert string columns to integers first var1 = group['Variable1'].astype(int) var2 = group['Variable2'].astype(int) return (var2 - var1).sum() # Apply the custom function to each group diff_results = subject_groups.apply(calculate_total_difference) print(diff_results)
4. Fill missing values within groups
Your Category column has blank entries. We can fill these using the most common value (mode) from the same subject's group:
# Fill blank Category values with the group's mode df['Category'] = subject_groups['Category'].transform( lambda x: x.fillna(x.mode()[0] if not x.mode().empty else '') ) print(df)
Pro Tip: Keep Subject as a Column (Not Index)
By default, groupby() makes Subject the index of the resulting object. If you want to keep it as a regular column, add as_index=False to the groupby call:
subject_groups_with_col = df.groupby('Subject', as_index=False)
内容的提问来源于stack exchange,提问作者Prometheus

