You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中拆分年份列完成数据集重构

首先明确你的原始数据中GDP列的格式为「年份: 增长率」,以下是两种可直接复用的R实现方案:

方法1:tidyverse 方案(代码简洁易维护)

先安装加载依赖包:

# 未安装的话先执行 install.packages("tidyverse")
library(tidyverse)

执行数据转换,假设你的原始数据框命名为df:

new_df <- df %>%
  # 拆分出年份和增长率列
  separate(GDP, into = c("年份", "GDP增长率"), sep = ":", convert = TRUE) %>%
  # 清理增长率列的多余空格
  mutate(GDP增长率 = str_trim(GDP增长率))

你也可以用字符串提取的方式实现相同效果:

new_df <- df %>%
  mutate(
    年份 = str_extract(GDP, "\\d{4}") %>% as.numeric(),
    GDP增长率 = str_trim(str_remove(GDP, "^\\d{4}:"))
  ) %>%
  # 可选:删除原始GDP列
  select(-GDP)

方法2:基础R方案(无需额外安装依赖)

# 拆分GDP列内容
split_content <- strsplit(df$GDP, ": ")
# 生成新列
df$年份 <- as.numeric(sapply(split_content, "[", 1))
df$GDP增长率 <- sapply(split_content, "[", 2)
# 可选:删除原始GDP列
df$GDP <- NULL

你可以用下方的样例测试数据验证代码效果:

# 测试用样例数据
df <- data.frame(
  地区 = c("北京", "上海", "广东"),
  GDP = c("2022: 0.7%", "2022: -0.2%", "2022: 1.9%")
)

如果你的GDP列格式和上述假设略有差异,调整分隔符或正则匹配规则即可适配。

内容的提问来源于stack exchange,提问作者Egor

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 02:48:01