You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言批量将教育机构名称缩写转换为官方全称的高效实现方法

教育机构缩写批量转全称的高效R实现方案

你现有一批教育机构缩写需要转换为全称,手动搜索录入效率过低,可通过以下两类方案实现批量处理:

你提供的可复现示例数据如下:

data <- structure(list(Affiliations = c("UNIV MELBOURNE", "UNIV NEWCASTLE", 
                                        "FORDHAM UNIV", "PRINCETON UNIV", 
                                        "CITY UNIV LONDON", "UNIV CONNECTICUT", 
                                        "EMORY UNIV", "NATL BUR ECON RES", 
                                        "NATL CHENGCHI UNIV", "OHIO STATE UNIV")), 
                  row.names = c(NA, -10L), 
                  class = c("tbl_df", "tbl", "data.frame"))

你当前使用的手动搜索录入代码如下:

library(searcher) # 调用浏览器搜索的功能包

df$new <- NA

for (i in 1:length(df$Affiliations)) {
  search_startpage(data$Affiliations[i], rlang = F)
  data$new[i] <- readline()
}

方案1:预构建缩写映射+匹配公共学术机构名录(最优方案,无反爬风险)

这是处理千条级别数据效率最高的方案,步骤如下:

  • 第一步:先整理通用缩写映射表,批量替换基础缩写:
    常见的机构缩写对应关系可提前整理,比如UNIV→University、NATL→National、BUR→Bureau、RES→Research,直接批量替换所有缩写字符串
  • 第二步:匹配公开的全球大学/研究机构名录,R环境下可直接调用openalexR、ror等工具包内置的机构元数据,直接匹配替换后的字符串得到标准全称
    示例代码:
# 1. 构建基础缩写映射
abbr_map <- c(
  "UNIV" = "University",
  "NATL" = "National",
  "BUR" = "Bureau",
  "RES" = "Research"
)
# 批量替换缩写
data$affil_fixed <- stringr::str_replace_all(data$Affiliations, abbr_map)

# 2. 调用openalexR匹配标准机构名
library(openalexR)
data$full_name <- purrr::map_chr(data$affil_fixed, ~{
  res <- oa_fetch(entities = "institutions", search = .x, count_only = F, limit = 1)
  ifelse(nrow(res) == 0, NA, res$display_name)
})

该方案可以覆盖90%以上的常规高校、研究机构,剩余匹配失败的再做人工核对即可。

方案2:使用rvest批量爬取搜索结果(备用方案)

如果特殊机构较多,无法通过名录匹配,可通过rvest批量爬取搜索引擎的首条结果提取全称,注意要设置请求间隔避免被封IP:
示例代码:

library(rvest)
library(httr)

data$full_name <- purrr::map_chr(data$Affiliations, function(abbr){
  # 设置请求头模拟浏览器
  headers <- c(
    "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
  )
  # 构造搜索链接,此处用必应搜索,无特殊反爬限制
  search_url <- paste0("https://cn.bing.com/search?q=", URLencode(paste0(abbr, " 机构全称")))
  # 发送请求
  res <- GET(search_url, add_headers(.headers=headers))
  # 解析结果,提取首条结果标题
  full_name <- read_html(res) |> 
    html_element(".b_algo h2") |> 
    html_text2()
  # 间隔2秒避免反爬
  Sys.sleep(2)
  return(full_name)
})

注意事项:爬取结果需要做二次校验,避免搜索结果出错。

内容的提问来源于stack exchange,提问作者Mohanasundaram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 05:18:02