You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest多页爬取欧洲可再生能源企业链接遇重复问题求助

问题分析与解决方案

你的代码仅返回第一页数据,核心原因是网站通过请求头识别爬虫,rvest::read_html()的默认请求标识会被拦截,所有page-*链接实际都返回第一页内容。

修复步骤

  • 自定义请求头,模拟浏览器访问
  • 用httr发送请求后再解析HTML,避免默认请求被拦截
  • 添加请求延迟,降低触发反爬机制的风险

修改后的代码

library(rvest)
library(httr)
library(purrr)

# 生成全部78页的链接
link <- paste0('https://www.energy-xprt.com/renewable-energy/companies/location-europe/page-', 1:78)

# 模拟浏览器请求头
headers <- add_headers(
  `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
  `Accept-Language` = "zh-CN,zh;q=0.9",
  `Referer` = "https://www.energy-xprt.com/renewable-energy/companies/location-europe/"
)

# 遍历链接并爬取数据
result <- map(link, function(x) {
  # 发送带请求头的请求
  response <- GET(x, headers)
  # 检查请求是否成功
  if (http_status(response)$category != "Success") {
    warning(paste("页面", x, "请求失败"))
    return(NULL)
  }
  # 解析HTML并提取企业链接
  content(response, as = "parsed") %>% 
    html_nodes("[class='h2 mb-0'] a") %>% 
    html_attr("href") %>% 
    unique()
  # 每次请求后延迟1秒,避免被封
  Sys.sleep(1)
}) %>% 
  unlist() %>% 
  unique()

关键说明

  1. 请求头伪装:通过User-Agent模拟真实浏览器,绕过网站的爬虫识别
  2. 请求状态校验:增加失败提示,方便排查异常页面
  3. 简化选择器:将html_nodes和html_elements合并为一个选择器,提升解析效率
  4. 延迟控制:Sys.sleep(1)降低请求频率,减少被IP封禁的概率

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 08:23:34