在R中基于区间扩展数据集行并替换标识后缀
解决方案
用R的tidyverse工具链可以高效实现需求,步骤如下:
- 提取
desc中的区间上下限:用正则表达式从desc里匹配出起始值和结束值 - 按区间生成数值序列,将每行数据展开为对应数量的行
- 把
first_id末尾的xx替换为序列中的数值
完整代码
# 安装并加载tidyverse(如果未安装) # install.packages("tidyverse") library(tidyverse) # 处理数据 result_df <- df %>% # 从desc提取起始和结束数值 mutate( start = as.numeric(str_extract(desc, "(?<=xx = )\\d+")), end = as.numeric(str_extract(desc, "(?<= à )\\d+")) ) %>% # 按区间生成数值序列,展开行 rowwise() %>% mutate(num = list(start:end)) %>% unnest(num) %>% # 替换first_id中的xx为对应数值 mutate(first_id = str_replace(first_id, "xx", as.character(num))) %>% # 保留需要的列(可选,根据需求调整) select(first_id, desc) # 查看结果 head(result_df, 25)
代码说明
str_extract配合正则表达式:(?<=xx = )\\d+匹配xx =后的数字作为起始值,(?<= à )\\d+匹配à后的数字作为结束值rowwise()+list(start:end)为每行生成对应的数值序列,unnest(num)将序列展开为多行str_replace把first_id里的xx替换成当前行的num值
运行后就能得到你需要的first_id序列,同时保留原desc内容。
内容的提问来源于stack exchange,提问作者Val
相关产品推荐
相关产品推荐

