You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Pandas对数据集中的个体进行分组

Grouping by Subject in Pandas: Step-by-Step Guide

Alright, let's walk through exactly how to group your dataset by the Subject column using Pandas. First, I noticed your Variable3 list was incomplete (only 5 elements instead of 12), so I'll fix that first to ensure we can work with a valid, runnable DataFrame.

Step 1: Fix and Load the Dataset

First, let's get our data into a proper Pandas DataFrame:

import pandas as pd

# Corrected full dataset (filled in missing Variable3 values)
d1 = {
    'Subject': ['Subject1','Subject1','Subject1','Subject2','Subject2','Subject2','Subject3','Subject3','Subject3','Subject4','Subject4','Subject4'], 
    'Event':['1','2','3','1','2','3','1','2','3','1','2','3'], 
    'Category':['1','1','2','2','1','2','2','','2','1','1',''], 
    'Variable1':['1','2','3','4','5','6','7','8','9','10','11','12'], 
    'Variable2':['12','11','10','9','8','7','6','5','4','3','2','1'], 
    'Variable3': ['-6','-5','-4','-3','-2','-1','0','1','2','3','4','5']
}
df = pd.DataFrame(d1)

Step 2: Basic Grouping by Subject

The core of grouping is Pandas' groupby() method. This creates a DataFrameGroupBy object that we can manipulate for all sorts of tasks:

# Group the DataFrame by the 'Subject' column
subject_groups = df.groupby('Subject')

Step 3: Common Operations on Grouped Data

Now that we have our groups, here are the most useful things you can do with them:

1. Iterate through groups to inspect data

If you want to see exactly what's in each subject's group, loop through the DataFrameGroupBy object:

for subject_name, group_data in subject_groups:
    print(f"--- Data for {subject_name} ---")
    print(group_data)
    print("\n")

2. Aggregate statistics (sum, mean, max, etc.)

Use agg() to compute multiple stats at once for each group. For example, let's calculate the mean and sum of Variable1, plus the max of Variable2 per subject:

# Aggregate multiple metrics
group_stats = subject_groups.agg(
    Variable1_average=('Variable1', 'mean'),
    Variable1_total=('Variable1', 'sum'),
    Variable2_highest=('Variable2', 'max')
)
print(group_stats)

Note: Since your Variable1/Variable2 are stored as strings, Pandas will convert them to numeric automatically for these calculations, but it's safer to explicitly cast them with astype(int) if you're working with larger datasets.

3. Apply custom functions to groups

For more complex logic, use apply() to run your own function on each group. Let's calculate the total difference between Variable2 and Variable1 for each subject:

def calculate_total_difference(group):
    # Convert string columns to integers first
    var1 = group['Variable1'].astype(int)
    var2 = group['Variable2'].astype(int)
    return (var2 - var1).sum()

# Apply the custom function to each group
diff_results = subject_groups.apply(calculate_total_difference)
print(diff_results)

4. Fill missing values within groups

Your Category column has blank entries. We can fill these using the most common value (mode) from the same subject's group:

# Fill blank Category values with the group's mode
df['Category'] = subject_groups['Category'].transform(
    lambda x: x.fillna(x.mode()[0] if not x.mode().empty else '')
)
print(df)

Pro Tip: Keep Subject as a Column (Not Index)

By default, groupby() makes Subject the index of the resulting object. If you want to keep it as a regular column, add as_index=False to the groupby call:

subject_groups_with_col = df.groupby('Subject', as_index=False)

内容的提问来源于stack exchange,提问作者Prometheus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:57:55