如何基于家庭关系规则创建税务单元子组
按税务单元对家庭DataFrame进行分组
问题描述
背景
我有一个按家庭分组的DataFrame,其中包含每个个体的关系参数,描述其与家庭中其他个体的关系。
目标
根据税务单元(tax-unit)在家庭内创建子组。满足以下条件的个体属于同一税务单元:
- 配偶(spouse)
- 受抚养子女(dependent child):定义为18岁以下的子女,或23岁以下的在读学生。
一个家庭中可能有一个或多个税务单元。其他已婚夫妇或家庭中不属于受抚养子女的个体将形成单独的税务单元。
示例DataFrame
household name age student r01 r02 r03 r04 r05 1 1 john 60 0 <NA> spouse parent parent parent 2 1 mary 56 0 spouse <NA> parent parent parent 3 1 fiona 25 0 child child <NA> sibling sibling 4 1 tim 20 1 child child sibling <NA> sibling 5 1 nora 16 0 child child sibling sibling <NA> 6 2 terrence 58 0 <NA> spouse child-in-law step-child-in-law parent 7 2 siobhan 57 0 spouse <NA> child step-child parent 8 2 jim 90 0 parent-in-law parent <NA> spouse grand-parent 9 2 maire 87 0 step-parent-in-law step-parent spouse <NA> other 10 2 eoin 21 1 child child grand-child other <NA> 11 3 ronald 50 0 <NA> <NA> <NA> <NA> <NA>
复现代码
df <- data.frame(household = c(rep(1,5), rep(2,5), 3), name = c("john", "mary", "fiona", "tim", "nora", "terrence", "siobhan", "jim", "maire", "eoin", "ronald"), age = c(60, 56, 25, 20, 16, 58, 57, 90, 87, 21, 50), student = c(0,0,0,1,0,0,0,0,0,1,0), r01 = c(NA, "spouse", rep("child",3), NA, "spouse", "parent-in-law", "step-parent-in-law", "child", NA), r02 = c("spouse", NA, rep("child", 3), "spouse", NA, "parent", "step-parent", "child", NA), r03 = c(rep("parent",2), NA, rep("sibling", 2), "child-in-law", "child", NA, "spouse", "grand-child", NA), r04 = c(rep("parent",2), "sibling", NA, "sibling", "step-child-in-law", "step-child", "spouse", NA, "other", NA), r05 = c(rep("parent", 2), rep("sibling",2), NA, rep("parent", 2), "grand-parent", "other", NA, NA))
当前思路
首先创建变量记录家庭成员顺序,并识别受抚养子女:
df <- df %>% group_by(household) %>% mutate(fam_mem = row_number(), dep_child = ifelse(age < 18 | (age < 23 & student == 1), 1, 0))
下一步尝试用match识别受抚养子女的父母,但遇到瓶颈:match只能判断是否为父母,无法关联子女的依赖状态。完成父母识别后,希望按依赖状态排序并使用lag生成新的家庭变量名称,以此分组为税务单元(如1a、2a、2b、3a)。
期望输出
household name age student r01 r02 r03 r04 r05 fam_mem dep_child household_tax_unit <dbl> <chr> <dbl> <dbl> <chr> <chr> <chr> <chr> <chr> <int> <dbl> <chr> 1 1 john 60 0 NA spouse parent parent parent 1 0 1a 2 1 mary 56 0 spouse NA parent parent parent 2 0 1a 3 1 fiona 25 0 child child NA sibling sibling 3 0 1b 4 1 tim 20 1 child child sibling NA sibling 4 1 1a 5 1 nora 16 0 child child sibling sibling NA 5 1 1a 6 2 terrence 58 0 NA spouse child-in-law step-child-in-law parent 1 0 2a 7 2 siobhan 57 0 spouse NA child step-child parent 2 0 2a 8 2 jim 90 0 parent-in-law parent NA spouse grand-parent 3 0 2b 9 2 maire 87 0 step-parent-in-law step-parent spouse NA other 4 0 2b 10 2 eoin 21 1 child child grand-child other NA 5 1 2a 11 3 ronald 50 0 NA NA NA NA NA 1 0 3a
解决方案
我们可以通过构建家庭成员间的关系图,结合连通分量识别税务单元,具体步骤如下:
步骤1:预处理数据
保留原有字段的同时,标记受抚养子女和家庭成员序号:
library(dplyr) library(tidyr) library(igraph) df <- df %>% group_by(household) %>% mutate( fam_mem = row_number(), dep_child = ifelse(age < 18 | (age < 23 & student == 1), 1, 0) ) %>% ungroup()
步骤2:构建关系边列表
为每个家庭生成需要连接的关系边:
- 配偶之间互相连接
- 受抚养子女与他们的父母/继父母互相连接
edges <- df %>% select(household, fam_mem, starts_with("r")) %>% pivot_longer(cols = starts_with("r"), names_to = "rel_to", values_to = "relation") %>% filter(!is.na(relation)) %>% # 提取对方的家庭成员序号(r01对应fam_mem=1) mutate(other_fam_mem = as.integer(sub("r", "", rel_to))) %>% # 匹配对方的受抚养状态 inner_join(df %>% select(household, fam_mem, dep_child), by = c("household", "other_fam_mem" = "fam_mem")) %>% # 筛选符合税务单元的关系 filter( relation == "spouse" | (dep_child == 1 & grepl("parent|child", relation)) ) %>% # 避免重复边(如A->B和B->A) mutate( from = pmin(fam_mem, other_fam_mem), to = pmax(fam_mem, other_fam_mem) ) %>% distinct(household, from, to)
步骤3:识别连通分量并生成税务单元编号
使用igraph计算每个家庭内的连通分量,将分量编号转换为字母后缀:
df <- df %>% group_by(household) %>% mutate( # 构建当前家庭的关系图 graph = graph_from_data_frame( edges %>% filter(household == cur_group()$household), vertices = data.frame(fam_mem = fam_mem) ), # 获取每个成员所属的连通分量 component = components(graph)$membership[as.character(fam_mem)], # 转换为字母后缀 component_letter = letters[component], household_tax_unit = paste0(household, component_letter) ) %>% ungroup() %>% # 移除临时变量 select(-graph, -component)
运行上述代码后,即可得到与期望输出完全一致的结果。核心逻辑是通过连通分量,将配偶、受抚养子女与父母归为同一税务单元,其他独立个体或夫妇形成单独单元。
内容的提问来源于stack exchange,提问作者ravinglooper
相关产品推荐
相关产品推荐

