使用rvest的download.file下载PDF出现空白/损坏问题咨询
download.file() Fails to Download PDFs Properly This is a super common issue, and it almost always boils down to using the wrong file mode for binary files like PDFs.
Here's the breakdown:
- By default,
download.file()usesmode = "w"(text mode) which is designed for plain text files. When you use this for binary files like PDFs, it messes up the file's byte structure—this is exactly why your downloaded PDF shows up blank or gets corrupted. - PDFs are binary data, so you need to use binary write mode (
mode = "wb") to preserve all the file's original bytes intact.
Fix 1: Add mode = "wb" to your code
Modify your command to include the binary mode flag:
download.file("https://www4.stat.ncsu.edu/~reich/ABA/Derivations3.pdf", destfile = "d3.pdf", mode = "wb")
This should download the PDF correctly, and it’ll open without any blank pages or corruption errors. While Linux/macOS might sometimes work without this flag, it’s critical for Windows systems—and it’s a best practice to include it for cross-platform compatibility.
Fix 2: Use httr for more robust downloads
If you still run into issues (like hidden server redirects or finicky header handling), the httr package offers a more reliable approach for grabbing binary files:
library(httr) # Fetch the PDF content as raw binary data pdf_response <- GET("https://www4.stat.ncsu.edu/~reich/ABA/Derivations3.pdf") # Write the raw data directly to a file writeBin(content(pdf_response, "raw"), "d3.pdf")
This method handles HTTP requests more explicitly and ensures the binary data is saved exactly as it’s received from the server.
内容的提问来源于stack exchange,提问作者Sean

