You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何规整县行政长官就职年份数据集?处理R中year字段格式问题

提取县行政长官的起始就职年份(R数据清洗方案)

问题背景

现有一份县行政长官数据集,year列格式混乱,包含纯年份、区间字符串(如"from 2001 to 2002")、日期字符串(如"01-feb-2003"),需要统一提取每位长官的起始就职年份,规整为数值型。

原始数据集:

df <- data.frame(year= c(2000, "from 2001 to 2002", "01-feb-2003", 2000, "01-jan-2002", "from 2004 to 2005"),
                  executive.name= c("Johnson", "Smith", "Alleghany", "Roberts", "Clarke", "Tollson"),
                  district= rep(c(1001, 1002), each=3))

解决方案

使用tidyverse工具链中的stringr处理字符串,快速提取起始年份:

# 加载依赖包
library(tidyverse)

df_clean <- df %>%
  # 统一将year列转为字符型,处理混合格式
  mutate(year = as.character(year)) %>%
  # 提取第一个出现的4位数字(即起始年份),转为数值型
  mutate(year = as.numeric(str_extract(year, "\\d{4}")))

结果验证

处理后的数据与预期一致:

> df_clean
  year executive.name district
1 2000         Johnson     1001
2 2001           Smith     1001
3 2003        Alleghany     1001
4 2000         Roberts     1002
5 2002          Clarke     1002
6 2004          Tollson     1002

代码逻辑说明

  • as.character(year):将原始混合类型的year列统一转为字符型,避免格式冲突
  • str_extract(year, "\\d{4}"):用正则表达式匹配所有4位数字的字符串,提取第一个匹配结果(对应起始年份)
  • as.numeric(...):将提取到的年份字符串转为数值型,保证数据类型统一

内容的提问来源于stack exchange,提问作者YouLocalRUser

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 01:11:08