在R中拆分含括号变体的型号列并扩展为独立多行数据
tidyverse 实现方案
library(tidyverse) result <- prices %>% rowwise() %>% mutate( # 提取所有括号内的可选值,自动去除前缀短横线 variants = str_extract_all(model, "\\(\\-?(.*?)\\)", group = 1) %>% unlist(), # 提取括号前的基准可选值 base_opt = str_extract(model, "([^-()]+)(?=\\()"), # 合并所有可选值 all_opts = c(base_opt, variants), # 生成型号模板,把可变段替换为占位符 template = str_replace(model, "[^-()]+(\\(.*?\\))+", "%s"), # 批量生成所有完整型号 model = sprintf(template, all_opts) ) %>% unnest_longer(model) %>% select(model, price)
执行后输出的result和你要求的desired_prices完全一致,自动适配:
- 可变部分单字符/多字符场景
- 可变部分在型号中间/末尾场景
- 单括号/多括号多变体场景
Base R 实现方案
如果你不想加载tidyverse依赖,可以用原生R代码实现:
result <- do.call(rbind, lapply(1:nrow(prices), function(i) { mod <- prices$model[i] prc <- prices$price[i] variants <- gsub("^\\-?|\\(|\\)", "", regmatches(mod, gregexpr("\\((.*?)\\)", mod))[[1]]) base_opt <- regmatches(mod, regexpr("([^-()]+)(?=\\()", mod, perl = TRUE)) all_opts <- c(base_opt, variants) template <- gsub("[^-()]+(\\(.*?\\))+", "%s", mod) data.frame(model = sprintf(template, all_opts), price = prc) }))
内容的提问来源于stack exchange,提问作者Rob Creel
相关产品推荐
相关产品推荐

