You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在dplyr中使用正则分组提取文献元数据报错,求修正方法

问题描述

我有如下字符串:

txt <- "Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22."

我需要从该字符串中提取作者姓名、年份和标题,在regex101上测试的正则命令可以正常运行:

result <- regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt))

result[[1]][2]
[1] "Harris P R, Harris D L"

result[[1]][3]
[1] "1983"

result[[1]][4]
[1] "Training for the Metaindustrial Work Culture"

现在我有一个包含多个此类字符串的数据框:

df <- data.frame(txt = c("Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22.",
"Cruise M J, Gorenberg B D (1985). The tools of management: keeping high touch in a high tech world. International nursing review, 32(6): 166-169, 173."))

我尝试用dplyr结合正则分组提取信息,编写了如下代码:

new_df <- df %>%
    rownames_to_column(var = "row_id") %>%
    mutate(result = regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt)),
           authors = result[[row_id]][2],
           year = result[[row_id]][3],
           title = result[[row_id]][4])

但运行时报错:

Error in `mutate()`:
! Problem while computing `authors = result[[row_id]][2]`.
Caused by error in `result[[row_id]]`:
! no such index at level 1
Run `rlang::last_error()` to see where the error occurred.

rlang::last_error()

<error/dplyr:::mutate_error>
Error in `mutate()`:
! Problem while computing `authors = result[[row_id]][2]`.
Caused by error in `result[[row_id]]`:
! no such index at level 1
---
Backtrace:
 1. df %>% rownames_to_column(var = "row_id") %>% ...
 3. dplyr:::mutate.data.frame(...)
 4. dplyr:::mutate_cols(.data, dplyr_quosures(...), caller_env = caller_env())
 6. mask$eval_all_mutate(quo)
Run `rlang::last_trace()` to see the full context.

请问需要修改哪些地方才能正常实现需求?


错误原因与修复方案

错误原因

regmatches处理向量输入时,返回的是嵌套列表(每个元素对应原数据框的一行)。而你在mutate中用result[[row_id]]的写法是错误的——row_id是数据框的列,不能直接用来索引整个result列表,应该针对每行的单个列表元素提取分组内容。

修复方案

有两种简洁高效的实现方式:

方案一:用purrr::map逐个提取列表元素

利用purrr的映射函数,对每行的匹配结果列表单独提取分组内容:

library(dplyr)
library(purrr)

new_df <- df %>%
  mutate(
    # 生成每行的匹配结果列表
    result = regmatches(txt, regexec("([^\\(]+) \\((\\d+)\\). ([^\\.]+).", txt)),
    # 用map提取每个分组,map_chr确保返回字符向量
    authors = map_chr(result, ~ .x[[2]]),
    year = map_chr(result, ~ .x[[3]]),
    title = map_chr(result, ~ .x[[4]])
  ) %>%
  # 可选:移除中间的result列
  select(-result)

方案二:用stringr::str_match直接生成矩阵(更简洁)

stringr::str_match可以直接对向量执行正则匹配,返回包含所有分组的矩阵,无需处理嵌套列表:

library(dplyr)
library(stringr)

new_df <- df %>%
  mutate(
    # str_match返回矩阵,第1列是完整匹配,第2-4列是目标分组
    match_matrix = str_match(txt, "([^\\(]+) \\((\\d+)\\). ([^\\.]+)\\."),
    authors = match_matrix[, 2],
    year = match_matrix[, 3],
    title = match_matrix[, 4]
  ) %>%
  select(-match_matrix)

验证结果

两种方案运行后,new_df都会得到如下结果:

txt                authors year                                                                   title
1 Harris P R, Harris D L (1983). Training for the Metaindustrial Work Culture. Journal of European Industrial Training, 7(7): 22. Harris P R, Harris D L 1983                     Training for the Metaindustrial Work Culture
2 Cruise M J, Gorenberg B D (1985). The tools of management: keeping high touch in a high tech world. International nursing review, 32(6): 166-169, 173.    Cruise M J, Gorenberg B D 1985 The tools of management: keeping high touch in a high tech world

内容的提问来源于stack exchange,提问作者aterhorst

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 23:55:20