You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于家庭关系规则创建税务单元子组

按税务单元对家庭DataFrame进行分组

问题描述

背景

我有一个按家庭分组的DataFrame,其中包含每个个体的关系参数,描述其与家庭中其他个体的关系。

目标

根据税务单元(tax-unit)在家庭内创建子组。满足以下条件的个体属于同一税务单元:

  • 配偶(spouse)
  • 受抚养子女(dependent child):定义为18岁以下的子女,或23岁以下的在读学生。

一个家庭中可能有一个或多个税务单元。其他已婚夫妇或家庭中不属于受抚养子女的个体将形成单独的税务单元。

示例DataFrame

household     name age student                r01         r02          r03               r04          r05
1          1     john  60       0               <NA>      spouse       parent            parent       parent
2          1     mary  56       0             spouse        <NA>       parent            parent       parent
3          1    fiona  25       0              child       child         <NA>           sibling      sibling
4          1      tim  20       1              child       child      sibling              <NA>      sibling
5          1     nora  16       0              child       child      sibling           sibling         <NA>
6          2 terrence  58       0               <NA>      spouse child-in-law step-child-in-law       parent
7          2  siobhan  57       0             spouse        <NA>        child        step-child       parent
8          2      jim  90       0      parent-in-law      parent         <NA>            spouse grand-parent
9          2    maire  87       0 step-parent-in-law step-parent       spouse              <NA>        other
10         2     eoin  21       1              child       child  grand-child             other         <NA>
11         3   ronald  50       0               <NA>        <NA>         <NA>              <NA>         <NA>

复现代码

df <- data.frame(household = c(rep(1,5), rep(2,5), 3),
           name = c("john", "mary", "fiona", "tim", "nora", "terrence", "siobhan", "jim", "maire", "eoin", "ronald"),
           age = c(60, 56, 25, 20, 16, 58, 57, 90, 87, 21, 50),
           student = c(0,0,0,1,0,0,0,0,0,1,0),
           r01 = c(NA, "spouse", rep("child",3), NA, "spouse", "parent-in-law", "step-parent-in-law", "child", NA),
           r02 = c("spouse", NA, rep("child", 3), "spouse", NA, "parent", "step-parent", "child", NA),
           r03 = c(rep("parent",2), NA, rep("sibling", 2), "child-in-law", "child", NA, "spouse", "grand-child", NA),
           r04 = c(rep("parent",2), "sibling", NA, "sibling", "step-child-in-law", "step-child", "spouse", NA, "other", NA),
           r05 = c(rep("parent", 2), rep("sibling",2), NA, rep("parent", 2), "grand-parent", "other", NA, NA))

当前思路

首先创建变量记录家庭成员顺序,并识别受抚养子女:

df <- df %>%
  group_by(household) %>%
  mutate(fam_mem = row_number(),
         dep_child = ifelse(age < 18 | (age < 23 & student == 1), 1, 0))

下一步尝试用match识别受抚养子女的父母,但遇到瓶颈:match只能判断是否为父母,无法关联子女的依赖状态。完成父母识别后,希望按依赖状态排序并使用lag生成新的家庭变量名称,以此分组为税务单元(如1a、2a、2b、3a)。

期望输出

household name       age student r01                r02         r03          r04               r05          fam_mem dep_child household_tax_unit
       <dbl> <chr>    <dbl>   <dbl> <chr>              <chr>       <chr>        <chr>             <chr>          <int>     <dbl> <chr>             
 1         1 john        60       0 NA                 spouse      parent       parent            parent             1         0 1a                
 2         1 mary        56       0 spouse             NA          parent       parent            parent             2         0 1a                
 3         1 fiona       25       0 child              child       NA           sibling           sibling            3         0 1b                
 4         1 tim         20       1 child              child       sibling      NA                sibling            4         1 1a                
 5         1 nora        16       0 child              child       sibling      sibling           NA                 5         1 1a                
 6         2 terrence    58       0 NA                 spouse      child-in-law step-child-in-law parent             1         0 2a                
 7         2 siobhan     57       0 spouse             NA          child        step-child        parent             2         0 2a                
 8         2 jim         90       0 parent-in-law      parent      NA           spouse            grand-parent       3         0 2b                
 9         2 maire       87       0 step-parent-in-law step-parent spouse       NA                other              4         0 2b                
10         2 eoin        21       1 child              child       grand-child  other             NA                 5         1 2a                
11         3 ronald      50       0 NA                 NA          NA           NA                NA                 1         0 3a  

解决方案

我们可以通过构建家庭成员间的关系图,结合连通分量识别税务单元,具体步骤如下:

步骤1:预处理数据

保留原有字段的同时,标记受抚养子女和家庭成员序号:

library(dplyr)
library(tidyr)
library(igraph)

df <- df %>%
  group_by(household) %>%
  mutate(
    fam_mem = row_number(),
    dep_child = ifelse(age < 18 | (age < 23 & student == 1), 1, 0)
  ) %>%
  ungroup()

步骤2:构建关系边列表

为每个家庭生成需要连接的关系边:

  • 配偶之间互相连接
  • 受抚养子女与他们的父母/继父母互相连接
edges <- df %>%
  select(household, fam_mem, starts_with("r")) %>%
  pivot_longer(cols = starts_with("r"), names_to = "rel_to", values_to = "relation") %>%
  filter(!is.na(relation)) %>%
  # 提取对方的家庭成员序号(r01对应fam_mem=1)
  mutate(other_fam_mem = as.integer(sub("r", "", rel_to))) %>%
  # 匹配对方的受抚养状态
  inner_join(df %>% select(household, fam_mem, dep_child), 
             by = c("household", "other_fam_mem" = "fam_mem")) %>%
  # 筛选符合税务单元的关系
  filter(
    relation == "spouse" | 
    (dep_child == 1 & grepl("parent|child", relation))
  ) %>%
  # 避免重复边(如A->B和B->A)
  mutate(
    from = pmin(fam_mem, other_fam_mem),
    to = pmax(fam_mem, other_fam_mem)
  ) %>%
  distinct(household, from, to)

步骤3:识别连通分量并生成税务单元编号

使用igraph计算每个家庭内的连通分量,将分量编号转换为字母后缀:

df <- df %>%
  group_by(household) %>%
  mutate(
    # 构建当前家庭的关系图
    graph = graph_from_data_frame(
      edges %>% filter(household == cur_group()$household),
      vertices = data.frame(fam_mem = fam_mem)
    ),
    # 获取每个成员所属的连通分量
    component = components(graph)$membership[as.character(fam_mem)],
    # 转换为字母后缀
    component_letter = letters[component],
    household_tax_unit = paste0(household, component_letter)
  ) %>%
  ungroup() %>%
  # 移除临时变量
  select(-graph, -component)

运行上述代码后,即可得到与期望输出完全一致的结果。核心逻辑是通过连通分量,将配偶、受抚养子女与父母归为同一税务单元,其他独立个体或夫妇形成单独单元。


内容的提问来源于stack exchange,提问作者ravinglooper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 07:55:55