R & dplyr:为组内选定成员分配组级别特征
Hey there! Let's work through how to assign group-level features to each individual in your dataset. From your description, you've got a large dataset split into contiguous groups (via grp), each with IDs starting at 1, plus is_child flags and momloc (0 or a group member's ID for the mother). Here are practical solutions using two common data analysis tools:
Using Python (Pandas)
1. Basic Group-Level Summary Features
If you want to assign aggregate stats like group size, proportion of children in the group, or average age to every individual in their group, this straightforward approach works:
import pandas as pd # Example dataset matching your structure data = pd.DataFrame({ 'grp': [1,1,1,1,2,2,2], 'id': [1,2,3,4,1,2,3], 'is_child': [False, True, True, False, True, False, True], 'momloc': [0, 1, 1, 0, 2, 0, 2], 'age': [30, 5, 3, 35, 4, 28, 2] }) # Calculate group-level stats group_stats = data.groupby('grp').agg( group_size=('id', 'count'), child_proportion=('is_child', 'mean'), avg_group_age=('age', 'mean') ).reset_index() # Merge stats back to every individual in the group data_with_group_features = pd.merge(data, group_stats, on='grp', how='left')
2. Mother-Derived Group-Level Features
If you need to pull attributes from a child's mother (using momloc), we'll handle this within each group since IDs are unique per group:
def add_mother_attributes(group): # Create a lookup map from group ID to individual features id_lookup = group.set_index('id')[['age', 'is_child']].to_dict('index') # Fetch mother's features where momloc isn't 0 group['mom_age'] = group['momloc'].apply( lambda x: id_lookup[x]['age'] if x != 0 else None ) group['mom_is_child'] = group['momloc'].apply( lambda x: id_lookup[x]['is_child'] if x != 0 else None ) return group # Apply the function to each group data_with_mom_features = data.groupby('grp', group_keys=False).apply(add_mother_attributes)
Using R (dplyr)
1. Basic Group-Level Summary Features
With dplyr, you can compute and assign group stats in one pipeline:
library(dplyr) # Example dataset data <- tibble( grp = c(1,1,1,1,2,2,2), id = c(1,2,3,4,1,2,3), is_child = c(FALSE, TRUE, TRUE, FALSE, TRUE, FALSE, TRUE), momloc = c(0, 1, 1, 0, 2, 0, 2), age = c(30,5,3,35,4,28,2) ) # Add group-level features to each row data_with_group_features <- data %>% group_by(grp) %>% mutate( group_size = n(), child_proportion = mean(is_child), avg_group_age = mean(age) ) %>% ungroup()
2. Mother-Derived Features
To pull a mother's attributes using momloc, we can match IDs within each group:
data_with_mom_features <- data %>% group_by(grp) %>% mutate( mom_age = case_when( momloc != 0 ~ age[match(momloc, id)], TRUE ~ NA_real_ ), mom_is_child = case_when( momloc != 0 ~ is_child[match(momloc, id)], TRUE ~ NA_logical_ ) ) %>% ungroup()
Quick Notes
- Since your groups are already contiguous, the grouping functions will work as expected. If you ever need to confirm, sort the data by
grpfirst (e.g.,data.sort_values('grp')in pandas,data %>% arrange(grp)in R). - Double-check that non-zero
momlocvalues always correspond to valid IDs in the same group. You can add a quick validation step (e.g., in pandas:assert (data.groupby('grp')['momloc'].apply(lambda x: x[x!=0].isin(x.id).all())).all()).
内容的提问来源于stack exchange,提问作者andrewH

