You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言的Google Scholar网页爬取迭代实现问题

解决Google Scholar批量爬取的迭代优化问题

你的核心问题是循环里缺少网页读取逻辑,且未定义存储网页的对象,导致wp找不到或者为NULL。以下是完整的迭代实现方案:

问题根源

  • 手动代码流程是「生成URL→读取网页→提取标题」,但你的循环直接跳过前两步,调用不存在的wp[[i]],自然报错。
  • Google Scholar每页最多10条结果,start参数按10递增(0,10,20...),需要批量生成对应URL。

完整迭代代码

# 加载所需包
packages <- c("rvest", "xml2", "data.table")
lapply(packages, library, character.only = TRUE)

# 配置爬取参数
author_name <- "pi campbell"  # 目标学者姓名
start_year <- 1998            # 起始年份
max_pages <- 4                # 要爬取的页数(对应手动版的4页)

# 生成所有爬取URL
start_values <- seq(from = 0, by = 10, length.out = max_pages)
urls <- sapply(start_values, function(x) {
  sprintf("http://scholar.google.com/scholar?start=%d&q=author:\"%s\"&as_ylo=%d",
          x, gsub(" ", "+", author_name), start_year)
})

# 初始化列表存储结果
titles_list <- vector("list", length(urls))

# 循环爬取
for (i in seq_along(urls)) {
  # 读取网页(加延迟避免反爬)
  Sys.sleep(2)
  wp <- tryCatch(xml2::read_html(urls[i]), 
                 error = function(e) {
                   message(sprintf("第%d页读取失败: %s", i, e$message))
                   return(NULL)
                 })
  
  # 提取标题(仅当网页读取成功时)
  if (!is.null(wp)) {
    titles_list[[i]] <- rvest::html_text(rvest::html_nodes(wp, '.gs_rt'))
  } else {
    titles_list[[i]] <- NA
  }
}

# 合并所有标题并去除NA
all_titles <- unlist(titles_list[!sapply(titles_list, is.na)])

代码说明

  1. 批量生成URL:用sprintf和sapply自动生成带不同start参数的URL,避免手动重复编写。
  2. 反爬处理:加入Sys.sleep(2),每次请求后暂停2秒,降低被Google封锁的风险。
  3. 错误处理:用tryCatch捕获网页读取失败的情况,避免程序直接崩溃,同时打印错误信息。
  4. 结果合并:最后把所有页面的标题合并成一个向量,方便后续处理。

进阶优化(自动爬取所有结果)

如果不确定学者有多少页成果,可以先爬第一页提取总结果数,再计算需要爬的页数:

# 先爬第一页获取总结果数
first_url <- sprintf("http://scholar.google.com/scholar?q=author:\"%s\"&as_ylo=%d",
                     gsub(" ", "+", author_name), start_year)
first_wp <- xml2::read_html(first_url)
total_results <- rvest::html_text(rvest::html_node(first_wp, ".gs_ab_mdw"))
total_results <- as.numeric(gsub("[^0-9]", "", total_results))
max_pages <- ceiling(total_results / 10)

# 后续爬取逻辑同上,替换max_pages为计算出的值即可

内容的提问来源于stack exchange,提问作者dbartram

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 02:20:00