如何在R中用case_when处理不同字符长度的fecha_nacimiento字段拆分
解决日期字符串拆分问题
先指出你原代码的几个错误:
case_when的括号配对逻辑混乱,每个条件的判断与对应表达式需要正确分隔- 在dplyr的
mutate操作里,无需用fecha_nacimiento$fecha_nacimiento调用列,直接写列名即可 nchar()的括号位置错误,应该是nchar(fecha_nacimiento) == 7,不能把判断逻辑放进nchar()的参数里
下面提供两种可行的实现方式:
方法一:补前导零统一格式后拆分
先把7位的日期字符串补一个前导零,统一为DDMMYYYY格式,之后就可以按固定位置拆分,不用分情况判断:
library(dplyr) # 构造示例数据 df <- tibble( fecha_nacimiento = c("1011997", "31122002"), n = c(1, 2) ) df_processed <- df %>% # 补前导零,将所有字符串统一为8位 mutate(fecha_8 = stringr::str_pad(fecha_nacimiento, width = 8, side = "left", pad = "0")) %>% # 按固定位置拆分日、月、年 mutate( days = substr(fecha_8, 1, 2), month = substr(fecha_8, 3, 4), year = substr(fecha_8, 5, 8) ) %>% # 移除临时生成的8位日期列(可选操作) select(-fecha_8) print(df_processed)
方法二:用case_when分别处理两种长度的字符串
如果不想补零,直接针对7位和8位的字符串分别指定拆分位置:
df_processed <- df %>% mutate( days = case_when( nchar(fecha_nacimiento) == 7 ~ substr(fecha_nacimiento, 1, 1), nchar(fecha_nacimiento) == 8 ~ substr(fecha_nacimiento, 1, 2) ), month = case_when( nchar(fecha_nacimiento) == 7 ~ substr(fecha_nacimiento, 2, 3), nchar(fecha_nacimiento) == 8 ~ substr(fecha_nacimiento, 3, 4) ), year = case_when( nchar(fecha_nacimiento) == 7 ~ substr(fecha_nacimiento, 4, 7), nchar(fecha_nacimiento) == 8 ~ substr(fecha_nacimiento, 5, 8) ) ) print(df_processed)
两种方法都能得到你需要的结果:
| fecha_nacimiento | n | days | month | year |
|---|---|---|---|---|
| 1011997 | 1 | 1 | 11 | 1997 |
| 31122002 | 2 | 31 | 12 | 2002 |
内容的提问来源于stack exchange,提问作者Facundo Javier Vargas
相关产品推荐
相关产品推荐

