在dplyr中使用正则分组提取文献元数据报错,求修正方法
问题描述
我有如下字符串:
txt <- "Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22."
我需要从该字符串中提取作者姓名、年份和标题,在regex101上测试的正则命令可以正常运行:
result <- regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt)) result[[1]][2] [1] "Harris P R, Harris D L" result[[1]][3] [1] "1983" result[[1]][4] [1] "Training for the Metaindustrial Work Culture"
现在我有一个包含多个此类字符串的数据框:
df <- data.frame(txt = c("Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22.", "Cruise M J, Gorenberg B D (1985). The tools of management: keeping high touch in a high tech world. International nursing review, 32(6): 166-169, 173."))
我尝试用dplyr结合正则分组提取信息,编写了如下代码:
new_df <- df %>% rownames_to_column(var = "row_id") %>% mutate(result = regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt)), authors = result[[row_id]][2], year = result[[row_id]][3], title = result[[row_id]][4])
但运行时报错:
Error in `mutate()`: ! Problem while computing `authors = result[[row_id]][2]`. Caused by error in `result[[row_id]]`: ! no such index at level 1 Run `rlang::last_error()` to see where the error occurred. rlang::last_error() <error/dplyr:::mutate_error> Error in `mutate()`: ! Problem while computing `authors = result[[row_id]][2]`. Caused by error in `result[[row_id]]`: ! no such index at level 1 --- Backtrace: 1. df %>% rownames_to_column(var = "row_id") %>% ... 3. dplyr:::mutate.data.frame(...) 4. dplyr:::mutate_cols(.data, dplyr_quosures(...), caller_env = caller_env()) 6. mask$eval_all_mutate(quo) Run `rlang::last_trace()` to see the full context.
请问需要修改哪些地方才能正常实现需求?
错误原因与修复方案
错误原因
regmatches处理向量输入时,返回的是嵌套列表(每个元素对应原数据框的一行)。而你在mutate中用result[[row_id]]的写法是错误的——row_id是数据框的列,不能直接用来索引整个result列表,应该针对每行的单个列表元素提取分组内容。
修复方案
有两种简洁高效的实现方式:
方案一:用purrr::map逐个提取列表元素
利用purrr的映射函数,对每行的匹配结果列表单独提取分组内容:
library(dplyr) library(purrr) new_df <- df %>% mutate( # 生成每行的匹配结果列表 result = regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt)), # 用map提取每个分组,map_chr确保返回字符向量 authors = map_chr(result, ~ .x[[2]]), year = map_chr(result, ~ .x[[3]]), title = map_chr(result, ~ .x[[4]]) ) %>% # 可选:移除中间的result列 select(-result)
方案二:用stringr::str_match直接生成矩阵(更简洁)
stringr::str_match可以直接对向量执行正则匹配,返回包含所有分组的矩阵,无需处理嵌套列表:
library(dplyr) library(stringr) new_df <- df %>% mutate( # str_match返回矩阵,第1列是完整匹配,第2-4列是目标分组 match_matrix = str_match(txt, "([^\\(]+) \\((\\d+)\\). ([^\\.]+)\\."), authors = match_matrix[, 2], year = match_matrix[, 3], title = match_matrix[, 4] ) %>% select(-match_matrix)
验证结果
两种方案运行后,new_df都会得到如下结果:
txt authors year title 1 Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22. Harris P R, Harris D L 1983 Training for the Metaindustrial Work Culture 2 Cruise M J, Gorenberg B D (1985). The tools of management: keeping high touch in a high tech world. International nursing review, 32(6): 166-169, 173. Cruise M J, Gorenberg B D 1985 The tools of management: keeping high touch in a high tech world
内容的提问来源于stack exchange,提问作者aterhorst
相关产品推荐
相关产品推荐

