You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Pandas DataFrame按单列分组转换为多级索引

Solution for Converting Pandas DataFrame to Indexed Style and Group Operations

Hey there! Let's break down how to achieve exactly what you want with your Pandas DataFrame, step by step.

Step 1: Replicate Your Original DataFrame

First, let's recreate the DataFrame you provided to work with:

import pandas as pd

data = {
    'group': [1, 1, 1, 2, 2, 3],
    'A': [1, 2, 4, 8, 5, 6],
    'B': [2, 3, 9, 1, 6, 5],
    'C': [3, 6, 9, 2, 4, 7]
}
df = pd.DataFrame(data)

Step 2: Convert to the Desired Display Format

To get the clean, repeated-index-hidden display you want, set group as the index and enable Pandas' sparse display option:

# Set 'group' as the DataFrame index
df_indexed = df.set_index('group')

# Turn on sparse display to hide duplicate index values in output
pd.set_option('display.multi_sparse', True)

Now when you print df_indexed, it will match your desired format perfectly:

A  B  C
group         
1      1  2  3
       2  3  6
       4  9  9
2      8  1  2
       5  6  4
3      6  5  7

Step 3: Access Sub-DataFrames by Group

You have a couple of flexible ways to pull up specific group data:

Option 1: Use .loc with Group Labels

If you want to directly access the sub-DataFrame for group 1, use index-based selection:

group_1_df = df_indexed.loc[1]

This returns all rows for group 1, and you can immediately run operations like mean() on it:

group_1_mean = group_1_df.mean()
# Output: A    2.333333, B    4.666667, C    6.000000

Option 2: Use groupby for Batch or Position-Based Access

For more control (like accessing groups by position, e.g., 0 for the first group), use groupby:

# Group the original DataFrame by 'group'
grouped = df.groupby('group')

# Get group 1's sub-DataFrame
group_1_df = grouped.get_group(1)

# Access the first group by position (matches group 1 in your data)
first_group_df = list(grouped)[0][1]

Step 4: Run Operations on Sub-DataFrames

All standard Pandas operations work seamlessly on these sub-DataFrames:

# Calculate mean for group 2
group_2_mean = grouped.get_group(2).mean()

# Calculate sum for group 3
group_3_sum = grouped.get_group(3).sum()

# Compute stats for all groups at once
all_groups_summary = grouped.agg(['mean', 'sum', 'max'])

内容的提问来源于stack exchange,提问作者AstroBen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:24:33