如何解决R语言download.file函数的403 Forbidden错误
解决download.file 403 Forbidden问题的几种方法
1. 补充完整请求头并指定请求方法
网站的反爬校验可能不止User-Agent,还会检查Referer、Cookie等字段。你可以打开浏览器访问目标PDF,在开发者工具的网络面板里复制完整请求头,把关键字段补充进去,同时指定method="libcurl"确保headers参数生效:
headers = c( `user-agent` = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.5005.61 Safari/537.36', `referer` = 'https://gibsons.civicweb.net/', `cookie` = '替换成你浏览器中该网站的Cookie内容' ) download.file("https://gibsons.civicweb.net/filepro/document/125764/Regular%20Council%20-%2006%20Dec%202022%20-%20Minutes%20-%20Pdf.pdf", "test.pdf", mode="wb", headers=headers, method="libcurl")
注意:Cookie可能会过期,需要定期从浏览器更新。
2. 改用httr包处理请求
download.file功能较基础,httr包能更灵活处理HTTP会话和请求配置:
library(httr) # 构造请求头 req_headers = add_headers( `user-agent` = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.5005.61 Safari/537.36', `referer` = 'https://gibsons.civicweb.net/' ) # 发送请求并保存文件 response = GET("https://gibsons.civicweb.net/filepro/document/125764/Regular%20Council%20-%2006%20Dec%202022%20-%20Minutes%20-%20Pdf.pdf", req_headers) writeBin(content(response, "raw"), "test.pdf")
如果网站需要会话验证,可以先访问首页建立会话再请求文件:
# 建立网站会话 session = html_session("https://gibsons.civicweb.net/") # 用会话请求目标文件 response = session %>% GET("https://gibsons.civicweb.net/filepro/document/125764/Regular%20Council%20-%2006%20Dec%202022%20-%20Minutes%20-%20Pdf.pdf") writeBin(content(response, "raw"), "test.pdf")
3. 确认URL有效性
部分网站的文件URL会附带临时token或时效参数,你爬取的URL可能已失效。可以重新从页面的下载按钮获取最新链接,或者模拟点击下载按钮的动作来抓取有效URL。
内容的提问来源于stack exchange,提问作者scotiaboy
相关产品推荐
相关产品推荐

