You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R爬取egamersworld电竞网站时遇403及数据获取难题

解决egamersworld.com爬虫403禁止访问及《使命召唤》赛事数据获取问题

可能触发403的反爬机制

  • 缺失浏览器专属请求头:现代浏览器会自动发送Sec-Fetch-*系列头(如Sec-Fetch-Site、Sec-Fetch-Mode),你的请求缺少这些字段,容易被识别为非浏览器请求。
  • Cookie验证缺失:网站可能要求请求携带首次访问时设置的Cookie,直接发起赛事页请求未携带Cookie会被拦截。
  • JS渲染依赖:如果赛事数据是通过JavaScript动态加载的,纯HTTP请求(如httr)无法获取渲染后的内容,同时这类网站常通过JS验证浏览器合法性。
  • WAF拦截:该网站可能使用了Cloudflare等防火墙工具,会对请求的合法性进行多维度校验,纯HTTP请求很难通过。

针对R的具体解决步骤

1. 补全完整浏览器请求头

把浏览器实际发送的关键请求头全部加入,模拟真实浏览器行为:

library(httr)

url <- "https://egamersworld.com/callofduty/matches"

headers <- add_headers(
  "Accept" = "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
  "Accept-Encoding" = "gzip, deflate",
  "Accept-Language" = "en-US,en;q=0.9",
  "Referer" = "https://egamersworld.com/matches",
  "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/133.0.0.0 Safari/537.36",
  "Sec-Fetch-Site" = "same-origin",
  "Sec-Fetch-Mode" = "navigate",
  "Sec-Fetch-Dest" = "document",
  "Upgrade-Insecure-Requests" = "1"
)

response <- GET(url, headers)
print(response$status_code)

2. 携带Cookie发起请求

先访问网站主页获取Cookie,再用该Cookie请求赛事页面:

library(httr)

# 先请求主页获取Cookie
home_response <- GET("https://egamersworld.com")
session_cookies <- cookies(home_response)

# 携带Cookie请求赛事页
match_url <- "https://egamersworld.com/callofduty/matches"
headers <- add_headers(
  "Accept" = "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8",
  "Accept-Encoding" = "gzip, deflate",
  "Accept-Language" = "en-US,en;q=0.9",
  "Referer" = "https://egamersworld.com/matches",
  "User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/133.0.0.0 Safari/537.36",
  "Sec-Fetch-Site" = "same-origin",
  "Sec-Fetch-Mode" = "navigate",
  "Sec-Fetch-Dest" = "document",
  "Upgrade-Insecure-Requests" = "1"
)

response <- GET(match_url, headers, set_cookies(.cookies = session_cookies))
print(response$status_code)

3. 模拟浏览器渲染(解决动态数据加载问题)

如果网站数据是JS动态生成的,使用chromote模拟无头浏览器获取渲染后的页面:

library(chromote)
library(rvest)

# 启动无头Chrome会话
browser_session <- ChromoteSession$new()
browser_session$Page$navigate("https://egamersworld.com/callofduty/matches")
browser_session$Page$waitForLoadEvent()

# 获取完整页面HTML
page_html <- browser_session$Runtime$evaluate("document.documentElement.outerHTML")$result$value
parsed_html <- read_html(page_html)

# 示例:提取赛事列表(需根据页面实际HTML结构调整选择器)
match_items <- parsed_html %>% html_elements(".match-card") %>% html_text2()
print(match_items)

# 关闭会话
browser_session$close()

关于JSON接口的说明

如果未找到公开的JSON数据接口,大概率是数据通过JS动态加载,或接口带有加密验证参数(如签名、token)。这种情况下,模拟浏览器渲染是最直接的解决方案,因为浏览器会自动处理JS加载和验证逻辑。

内容的提问来源于stack exchange,提问作者Nick Amato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 03:21:15