如何用dplyr找出每行首次出现0的列并提取列名后缀数字
Tidyverse解决方案:识别每行第一个含0的列并提取列名数字
给定如下数据集:
df <- data.frame(id=c(1:4), time_1=c(1, 0.9, 0.2, 0), time_2=c(0.1, 0.4, 0, 0.9), time_3=c(0,0.5,0.3,1.0))
数据集预览:
| id | time_1 | time_2 | time_3 |
|---|---|---|---|
| 1 | 1.0 | 0.1 | 0 |
| 2 | 0.9 | 0.4 | 0.5 |
| 3 | 0.2 | 0 | 0.3 |
| 4 | 0 | 0.9 | 1.0 |
需求:为每行识别第一个包含0的列,提取列名的最后一位数字生成count列(无0的行填NA),最终得到如下结果:
| id | time_1 | time_2 | time_3 | count |
|---|---|---|---|---|
| 1 | 1.0 | 0.1 | 0 | 3 |
| 2 | 0.9 | 0.4 | 0.5 | NA |
| 3 | 0.2 | 0 | 0.3 | 2 |
| 4 | 0 | 0.9 | 1.0 | 1 |
方法1:使用rowwise() + c_across()
library(tidyverse) df_result <- df %>% rowwise() %>% mutate( # 获取每行time列中第一个0的位置 first_zero_pos = which(c_across(starts_with("time_")) == 0)[1], # 提取对应列名的数字,无0则返回NA count = ifelse(is.na(first_zero_pos), NA_integer_, as.integer(str_extract(colnames(.)[first_zero_pos + 1], "\\d+"))) ) %>% ungroup() %>% select(id, time_1, time_2, time_3, count)
方法2:使用pmap_int()(更简洁)
library(tidyverse) df_result <- df %>% mutate( count = pmap_int(select(., starts_with("time_")), ~ { vals <- c(...) first_zero_pos <- which(vals == 0)[1] if (is.na(first_zero_pos)) NA_integer_ else as.integer(str_extract(names(vals)[first_zero_pos], "\\d+")) }) )
代码解释
select(., starts_with("time_")):筛选出所有以time_开头的列,作为逐行处理的对象pmap_int()/rowwise():对每行的time列值进行迭代处理,前者更高效,后者更直观which(vals == 0)[1]:找到当前行中第一个值为0的元素位置,无匹配时返回NAstr_extract(..., "\\d+"):提取对应列名中的数字部分,再转为整数类型
运行上述代码后,df_result的输出与需求完全一致:
> df_result # A tibble: 4 × 5 id time_1 time_2 time_3 count <int> <dbl> <dbl> <dbl> <int> 1 1 1 0.1 0 3 2 2 0.9 0.4 0.5 NA 3 3 0.2 0 0.3 2 4 4 0 0.9 1 1
内容的提问来源于stack exchange,提问作者jeff
相关产品推荐
相关产品推荐

