You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest的download.file下载PDF出现空白/损坏问题咨询

Why download.file() Fails to Download PDFs Properly

This is a super common issue, and it almost always boils down to using the wrong file mode for binary files like PDFs.

Here's the breakdown:

  • By default, download.file() uses mode = "w" (text mode) which is designed for plain text files. When you use this for binary files like PDFs, it messes up the file's byte structure—this is exactly why your downloaded PDF shows up blank or gets corrupted.
  • PDFs are binary data, so you need to use binary write mode (mode = "wb") to preserve all the file's original bytes intact.

Fix 1: Add mode = "wb" to your code

Modify your command to include the binary mode flag:

download.file("https://www4.stat.ncsu.edu/~reich/ABA/Derivations3.pdf", destfile = "d3.pdf", mode = "wb")

This should download the PDF correctly, and it’ll open without any blank pages or corruption errors. While Linux/macOS might sometimes work without this flag, it’s critical for Windows systems—and it’s a best practice to include it for cross-platform compatibility.

Fix 2: Use httr for more robust downloads

If you still run into issues (like hidden server redirects or finicky header handling), the httr package offers a more reliable approach for grabbing binary files:

library(httr)
# Fetch the PDF content as raw binary data
pdf_response <- GET("https://www4.stat.ncsu.edu/~reich/ABA/Derivations3.pdf")
# Write the raw data directly to a file
writeBin(content(pdf_response, "raw"), "d3.pdf")

This method handles HTTP requests more explicitly and ensures the binary data is saved exactly as it’s received from the server.

内容的提问来源于stack exchange,提问作者Sean

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:10:02