You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中从字符串提取多个employee_id并生成对应变量

在R中提取employee_id并生成单独变量

方法一:通过JSON解析(推荐,更稳定)

你的原始字符串是多个JSON对象的拼接,先转为标准JSON数组格式,再用专业工具解析提取:

  1. 安装并加载jsonlite包(R中处理JSON的常用工具)
# 若未安装则执行此句
install.packages("jsonlite")
library(jsonlite)
  1. 处理字符串并提取ID
# 原始字符串
raw_str <- '{"name":"John Smith","employee_id":"AB001","email":"johnsmith@abc.com"},{"name":"Sarah John","employee_id":"AB002","email":"sarahjohn@abc.com"},{"name":"Chris Lee","employee_id":"AB003","email":"chrislee@abc.com"},{"name":"Brian Johns","employee_id":"AB004","email":"brianjohns@abc.com"}'

# 包裹成标准JSON数组
json_str <- paste0("[", raw_str, "]")

# 解析JSON并提取所有employee_id
employee_ids <- fromJSON(json_str)$employee_id

# 转为目标格式的数据集,每个ID对应单独变量
result_df <- as.data.frame(t(employee_ids))
colnames(result_df) <- paste0("Employee ID", 1:length(employee_ids))

运行后result_df就是你需要的结构,打印结果如下:

Employee ID1 Employee ID2 Employee ID3 Employee ID4
1        AB001        AB002        AB003        AB004

方法二:正则表达式提取(无需额外包)

如果不想用JSON工具,可以直接用正则匹配提取ID值:

# 原始字符串
raw_str <- '{"name":"John Smith","employee_id":"AB001","email":"johnsmith@abc.com"},{"name":"Sarah John","employee_id":"AB002","email":"sarahjohn@abc.com"},{"name":"Chris Lee","employee_id":"AB003","email":"chrislee@abc.com"},{"name":"Brian Johns","employee_id":"AB004","email":"brianjohns@abc.com"}'

# 匹配并提取employee_id的内容
employee_ids <- regmatches(raw_str, gregexpr('"employee_id":"([^"]+)"', raw_str))[[1]]
employee_ids <- gsub('"employee_id":"([^"]+)"', '\\1', employee_ids)

# 转为目标格式数据集
result_df <- as.data.frame(t(employee_ids))
colnames(result_df) <- paste0("Employee ID", 1:length(employee_ids))

内容的提问来源于stack exchange,提问作者Brian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 05:15:07