如何清洗year列并保留列数据?实现年份值标准化处理
R数据框年份列清洗实现
原始数据
# 原始数据结构 df <- structure(list(year = c("Mar-10", "2014", "May-August", "2009/2010", "2015", NA_character_), date = c("August 31st, 2010", "March 13th, 2015", "May 31st, 2010", "June 16th, 2010", "May 18th, 2010", "April 7th, 2010")), row.names = c(NA, -6L), class = c("tbl_df", "tbl", "data.frame")) # 查看数据 df # # A tibble: 6 × 2 # year date # <chr> <chr> # 1 Mar-10 August 31st, 2010 # 2 2014 March 13th, 2015 # 3 May-August May 31st, 2010 # 4 2009/2010 June 16th, 2010 # 5 2015 May 18th, 2010 # 6 NA April 7th, 2010
清洗需求
需按照以下规则清洗year列,保留所有行:
- 若
year为标准4位年份(如"2014"),直接保留原值; - 若
year为短格式年月(如"Mar-10"),转换为对应4位年份(如2010); - 若
year无法确定单一整年(如"May-August"、"2009/2010")或为NA,从date列提取年份填充。
实现代码
使用dplyr结合stringr完成清洗,代码如下:
library(dplyr) library(stringr) cleaned_df <- df %>% # 从date列提取年份作为备用值 mutate(date_year = str_extract(date, "\\d{4}")) %>% # 按规则处理year列 mutate( year = case_when( # 匹配标准4位年份,直接保留 str_detect(year, "^\\d{4}$") ~ year, # 匹配MM-YY格式,转换为20YY的4位年份 str_detect(year, "^[A-Za-z]{3}-\\d{2}$") ~ paste0("20", str_extract(year, "\\d{2}")), # 其他所有情况,使用date列提取的年份 TRUE ~ date_year ) ) %>% # 移除临时辅助列 select(-date_year) # 查看清洗后的数据 cleaned_df # # A tibble: 6 × 2 # year date # <chr> <chr> # 1 2010 August 31st, 2010 # 2 2014 March 13th, 2015 # 3 2010 May 31st, 2010 # 4 2010 June 16th, 2010 # 5 2015 May 18th, 2010 # 6 2010 April 7th, 2010
代码解释
- 提取备用年份:用
str_extract(date, "\\d{4}")从date字符串中匹配4位数字,得到每个日期对应的年份,存入临时列date_year; - 分支处理year列:
- 第一个分支:通过正则
^\\d{4}$识别标准4位年份,直接保留; - 第二个分支:通过正则
^[A-Za-z]{3}-\\d{2}$识别"Mar-10"这类短格式年月,提取末尾两位数字并拼接"20"得到4位年份; - 第三个分支:覆盖所有无法直接识别的情况(范围型年份、NA等),直接使用
date_year的值;
- 第一个分支:通过正则
- 清理临时列:最后移除辅助用的
date_year,得到符合要求的结果。
内容的提问来源于stack exchange,提问作者DariusMT7
相关产品推荐
相关产品推荐

