You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用separate()拆分调查中特殊命名的年龄列

问题描述

我有一份精简后的调查结果数据,结构如下:

structure(list(`What is your age?` = c("65+", "65+", "65+", "25-34", "45-54", "65+"), `Gender identity` = c("Female", "Female", "Male", "Non-Binary", "Female", "Female")), row.names = 3:8, class = "data.frame")

想要把「What is your age?」列拆成「Min Age」和「Max Age」列:

  • 带范围的年龄(比如25-34)正常拆分
  • 65+的「Max Age」留空

试了几种separate()写法都报错:

workingfile$`What is your age?` %>% separate(`What is your age?`, c('Min Age', 'Max Age'), "_|(?=...$) ", convert = TRUE)
workingfile %>% separate(`What is your age?`, c('Min Age', 'Max Age'), "_|(?=...$) ", convert = TRUE)
workingfile %>% separate(.$`What is your age?`, c('Min Age', 'Max Age'), "_|(?=...$) ", convert = TRUE)

报错提示找不到What is your age?对象。

正确写法及说明

你的问题出在正则表达式错误和dplyr语法使用不当两个地方,修正后的代码如下:

library(dplyr)
library(tidyr)

# 先构造测试数据(如果已有workingfile可以跳过这行)
workingfile <- structure(list(`What is your age?` = c("65+", "65+", "65+", "25-34", "45-54", "65+"), `Gender identity` = c("Female", "Female", "Male", "Non-Binary", "Female", "Female")), row.names = 3:8, class = "data.frame")

# 执行拆分
result <- workingfile %>%
  separate(
    col = `What is your age?`,
    into = c("Min Age", "Max Age"),
    sep = "(-|\\+)",  # 匹配'-'或'+'作为分隔符,+需要转义
    convert = TRUE    # 自动把拆分结果转成数值,空值会变成NA
  )

关键说明:

  1. 正则修正:原来的正则"_|(?=...$) "完全不符合需求,换成"(-|\\+)"后,能精准匹配年龄范围的-和65+的+,实现正确拆分。
  2. 语法修正:dplyr管道中,separate()的第一个参数是数据框(管道自动传入),直接用col = 列名指定要拆分的列即可,不需要用.$调用。
  3. 自动处理空值:convert = TRUE会把65+拆分后的空字符串转为NA,刚好满足「Max Age留空」的需求,无需额外处理。

如果不需要保留原有的「What is your age?」列,默认remove = TRUE会自动删除,不用额外设置;要是想保留,加上remove = FALSE就行。

内容的提问来源于stack exchange,提问作者Plot Device

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 19:11:30