如何用Pandas在Jupyter Notebook中提取各家庭最高收入客户全量数据?
Extract Highest Income Member per Family in Pandas (Keep All Columns)
Got it, let's break down how to solve this problem effectively. Depending on whether you want only one member per family (the first one with max income) or all members who share the highest income in their family, here are the two common approaches:
Scenario 1: Get the first member with the highest income per family
If you just need one row per family (the first occurrence of the highest earner), use groupby() combined with idxmax() to get the indices of those rows, then slice your DataFrame with loc:
# Replace 'family_id' with your actual household identifier column # Replace 'income' with your income column name max_income_indices = df.groupby('family_id')['income'].idxmax() result = df.loc[max_income_indices]
Scenario 2: Keep all members who have the highest income in their family
If multiple family members have the same top income and you want to keep all of them, use transform() to compute the max income for each family, then filter rows where the income matches this family-level max:
# Calculate the maximum income for each family, broadcast to all rows in the family family_top_income = df.groupby('family_id')['income'].transform('max') # Filter rows where the member's income equals their family's top income result = df[df['income'] == family_top_income]
Key Notes:
- Make sure to replace
'family_id'and'income'with the actual column names from your DataFrame. - Both methods preserve all columns in your original table, which is exactly what you need.
内容的提问来源于stack exchange,提问作者Barbara Chen
相关产品推荐
相关产品推荐

