使用R爬取网站数据遇HTTP 403错误,添加User Agent仍未解决求指导
解决R爬取网站403 Forbidden的建议
补充完整请求头字段
仅设置User Agent可能不够,很多网站会校验Referer、Accept、Accept-Language等完整请求头。尝试补充这些字段:library(httr) library(rvest) headers <- add_headers( `User-Agent` = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36", `Referer` = "https://www.whosampled.com/", `Accept` = "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8", `Accept-Language` = "en-US,en;q=0.5" ) link <- GET("https://www.whosampled.com/Daft-Punk/Harder,-Better,-Faster,-Stronger/sampled/", headers) page <- read_html(link)添加请求延迟
频繁的请求会触发网站的反爬机制,在请求之间加入随机延迟:# 每次请求前随机等待2-5秒 Sys.sleep(runif(1, min=2, max=5)) link <- GET("https://www.whosampled.com/Daft-Punk/Harder,-Better,-Faster,-Stronger/sampled/", headers) page <- read_html(link)使用会话保持
模拟浏览器的会话状态,通过handle维持cookie等信息,有些网站会校验会话一致性:# 创建会话句柄 h <- handle("https://www.whosampled.com/") # 先访问首页获取cookie,再请求目标页面 GET("https://www.whosampled.com/", headers, handle = h) Sys.sleep(2) link <- GET("https://www.whosampled.com/Daft-Punk/Harder,-Better,-Faster,-Stronger/sampled/", headers, handle = h) page <- read_html(link)检查网站robots.txt
先确认目标路径是否被网站禁止爬取,访问https://www.whosampled.com/robots.txt查看Disallow列表,如果目标路径在其中,建议停止爬取或联系网站获取授权。尝试代理IP(备选)
如果你的IP被网站封禁,可以使用代理IP发送请求:link <- GET( url = "https://www.whosampled.com/Daft-Punk/Harder,-Better,-Faster,-Stronger/sampled/", headers, use_proxy(url = "你的代理IP", port = 你的代理端口) )
内容的提问来源于stack exchange,提问作者GiulioSurya
相关产品推荐
相关产品推荐

