如何按前缀条件批量处理tibble并编写通用函数?
处理带前缀的tibble列重编码问题
需求说明
- 对tibble每行执行如下规则:
- 若某前缀(如
a或b)的纯前缀列值为1,且该前缀下其余列均为0,则将该前缀下所有列重编码为1 - 若该前缀下任意列已为1,则保持该前缀下所有列不变
- 若某前缀(如
- 处理完成后删除所有纯前缀列(如
a、b列)
可复现示例数据
library(tibble) original_df <- tibble( a = c(1, 1, 0, 0, 1), a.1 = c(0, 0, 1, 0, 1), a.2 = c(0, 0, 0, 1, 0), b = c(0, 0, 0, 0, 1), b.1 = c(0, 0, 0, 1, 0), b.2 = c(0, 0, 0, 0, 0) ) original_df #> # A tibble: 5 × 6 #> a a.1 a.2 b b.1 b.2 #> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> #> 1 1 0 0 0 0 0 #> 2 1 0 0 0 0 0 #> 3 0 1 0 0 0 0 #> 4 0 0 1 0 1 0 #> 5 1 1 0 1 0 0
预期处理结果
expected_df <- tibble( a.1 = c(1, 0, 1, 0, 1), a.2 = c(1, 0, 0, 1, 0), b.1 = c(0, 0, 0, 1, 1), b.2 = c(0, 0, 0, 0, 1) ) expected_df #> # A tibble: 5 × 4 #> a.1 a.2 b.1 b.2 #> <dbl> <dbl> <dbl> <dbl> #> 1 1 1 0 0 #> 2 0 0 0 0 #> 3 1 0 0 0 #> 4 0 1 1 0 #> 5 1 0 1 1
通用解决方案函数
以下是适配任意前缀数量、任意前缀列数的通用函数:
library(dplyr) library(stringr) process_prefix_cols <- function(df) { # 提取所有纯前缀列(默认匹配纯字母列名,可根据需求修改正则) prefixes <- df %>% select(where(~ str_detect(colnames(df)[cur_column()], "^[a-zA-Z]+$"))) %>% colnames() # 遍历每个前缀处理 for (prefix in prefixes) { # 获取当前前缀的所有列(含纯前缀列和带后缀的列) prefix_cols <- str_subset(colnames(df), paste0("^", prefix, "(\\.|$)")) # 拆分纯前缀列和带后缀列 base_col <- prefix suffix_cols <- setdiff(prefix_cols, base_col) # 判断每行是否满足重编码条件:纯前缀列为1,且同前缀其余列全为0 need_recode <- df[[base_col]] == 1 & rowSums(df[suffix_cols]) == 0 # 对满足条件的行重编码 df <- df %>% mutate(across(all_of(suffix_cols), ~ ifelse(need_recode, 1, .x))) } # 删除所有纯前缀列,返回结果 df %>% select(-all_of(prefixes)) }
函数测试
将示例数据传入函数验证结果:
processed_df <- process_prefix_cols(original_df) processed_df #> # A tibble: 5 × 4 #> a.1 a.2 b.1 b.2 #> <dbl> <dbl> <dbl> <dbl> #> 1 1 1 0 0 #> 2 0 0 0 0 #> 3 1 0 0 0 #> 4 0 1 1 0 #> 5 1 0 1 1 # 验证与预期结果一致 all.equal(processed_df, expected_df) #> [1] TRUE
函数说明
- 自动识别纯字母列名为前缀,若你的前缀包含下划线、数字等,可修改
str_detect中的正则表达式(比如改成^[a-zA-Z0-9_]+$) - 对每个前缀分组单独判断处理,无需提前指定前缀数量或列数
- 仅对满足条件的行执行重编码,其余行保持原始值不变
内容的提问来源于stack exchange,提问作者T. C. Nobel
相关产品推荐
相关产品推荐

