You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言实现:从可变数量分隔的作者列提取指定作者

解决方案

我们可以通过拆分作者字符串、根据作者数量提取对应位置的作者来实现需求,以下是两种符合要求的R实现方式:

方法一:Base R 实现

# 原始数据
pub1 <- structure(list(publication = c("pub1", "pub2", "pub3", "pub4", 
        "pub5", "pub6"), authors = c("author1", "author1, author2", "author1, author2, author3", 
        "author1, author2, author3, author4", "author1, author2, author3, author4, author5", 
        "author1, author2, author3, author4, author5, author6")), 
        class = "data.frame", row.names = c(NA, -6L))

# 将作者字符串拆分为列表
author_list <- strsplit(pub1$authors, ", ")

# 提取第一作者
pub1$author_first <- sapply(author_list, function(x) x[1])

# 提取最后作者:作者数≥2时取最后一位并加前置空格,否则为空
pub1$author_last <- sapply(author_list, function(x) {
  if (length(x) >= 2) paste0(" ", x[length(x)]) else ""
})

# 提取倒数第二作者:作者数≥3时取倒数第二位并加前置空格,否则为空
pub1$author_second_last <- sapply(author_list, function(x) {
  if (length(x) >= 3) paste0(" ", x[length(x)-1]) else ""
})

# 查看最终结果
pub1

方法二:tidyverse 工具链实现

如果习惯使用管道式写法,可借助dplyr和purrr完成:

library(dplyr)
library(purrr)

pub1_processed <- pub1 %>%
  # 拆分作者字符串为列表列
  mutate(author_split = strsplit(authors, ", ")) %>%
  # 按规则提取对应作者
  mutate(
    author_first = map_chr(author_split, ~ .x[1]),
    author_last = map_chr(author_split, ~ if (length(.x) >=2) paste0(" ", .x[length(.x)]) else ""),
    author_second_last = map_chr(author_split, ~ if (length(.x) >=3) paste0(" ", .x[length(.x)-1]) else "")
  ) %>%
  # 移除临时生成的拆分列
  select(-author_split)

# 查看最终结果
pub1_processed

执行上述任意一种方法后,输出结果将与你给出的期望格式完全匹配:仅1位作者时仅保留第一作者,2位作者时提取第一和最后作者,3位及以上作者时提取第一、倒数第二和最后作者。

内容的提问来源于stack exchange,提问作者K. Wamae

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 19:30:42