基于ID分组与教育水平规则为DataFrame创建新列
Got it, let's work through this problem together. Since you're dealing with a pandas DataFrame, I'll walk you through a clear, practical solution—including filling in the missing logic for IDs with no graduated education records (I'll use a sensible default, but you can tweak it to fit your needs).
1. Sample DataFrame
First, let's define a sample dataset to make the logic concrete. I'll use a placeholder column is_graduated to flag completed education levels—swap this out for your actual graduation indicator (like a status column or completion date check).
import pandas as pd # Example data matching your description data = { 'id': [1, 1, 2, 2, 3, 3], 'education': [2, 4, 1, 3, 2, 5], 'is_graduated': [True, True, False, True, False, False] } df = pd.DataFrame(data)
2. Core Logic (Custom Function Approach)
This approach is straightforward and easy to modify if you need to adjust the graduation check:
def get_top_graduated_edu(group): # Filter rows where the education level was completed graduated_entries = group[group['is_graduated'] == True] if not graduated_entries.empty: # Return the highest completed education level for the ID return graduated_entries['education'].max() else: # Handle IDs with no graduated records (adjust this value as needed) return pd.NA # Apply the function to each ID group and add the new column df['highest_graduated_edu'] = df.groupby('id').apply(get_top_graduated_edu).reset_index(drop=True)
3. Faster Vectorized Alternative
For large datasets, this vectorized method is more efficient (avoids looping with custom functions):
# Keep education values only for graduated entries, set others to NaN df['graduated_edu'] = df['education'].where(df['is_graduated'] == True) # Compute max per ID (automatically ignores NaN for IDs with no graduates) df['highest_graduated_edu'] = df.groupby('id')['graduated_edu'].transform('max')
4. Customization Tips
- Adjust Graduation Check: Replace
df['is_graduated'] == Truewith your actual condition (e.g.,df['completion_status'] == 'Completed'). - Handle No-Graduate IDs: Change the
pd.NAin the custom function to whatever makes sense for your use case—like0,'No Completed Education', or the highest non-graduated level if that's what you need.
内容的提问来源于stack exchange,提问作者Lucas Dresl

