如何在R中下载文件时获取并使用浏览器默认的原始文件名?
在R中获取浏览器默认文件名并下载文件的解决方案
问题说明
需要下载链接 http://www.coingecko.com/price_charts/export/1/usd.xls 的文件,浏览器下载时自动命名为 btc-usd-max.xls,但直接使用 download.file(url=link, destfile=basename(link), method='auto') 会保存为 usd.xls。尝试过依赖 content-disposition 响应头的方案,但发现该字段未正确获取到。
可行解决方案
方法1:利用curl选项自动匹配浏览器文件名
借助curl工具的-J(遵循服务器建议的文件名)和-L(自动跟随重定向)参数,让download.file直接保存为浏览器使用的文件名:
link <- "http://www.coingecko.com/price_charts/export/1/usd.xls" download.file( url = link, destfile = tempfile(), # 临时占位,curl会自动替换为正确文件名 method = "curl", extra = "-J -L" )
注意:此方法要求系统已安装curl工具,Windows用户可通过WSL或单独安装curl环境实现。
方法2:通过httr追踪重定向提取文件名
服务器可能在重定向后的响应中携带文件名信息,可使用httr开启重定向追踪后提取:
library(httr) link <- "http://www.coingecko.com/price_charts/export/1/usd.xls" # 发送HEAD请求并跟随重定向 hd <- HEAD(link, config(followlocation = TRUE)) # 优先从响应头提取文件名,失败则从重定向后的URL获取 filename <- if (!is.null(headers(hd)$`content-disposition`)) { gsub(".*filename=\"?([^\"]+)\"?", "\\1", headers(hd)$`content-disposition`) } else { basename(hd$url) } # 用获取到的文件名下载文件 download.file(link, destfile = filename, method = "auto")
方法3:用curl包底层控制获取文件名
curl包提供更精细的请求控制,可直接解析响应头提取文件名:
library(curl) link <- "http://www.coingecko.com/price_charts/export/1/usd.xls" # 创建curl句柄并开启重定向 h <- new_handle(followlocation = TRUE) # 获取响应头信息 resp <- curl_fetch_memory(link, handle = h) # 解析响应头提取文件名 headers <- parse_headers(resp$headers) filename <- grep("content-disposition", headers, value = TRUE) filename <- if (length(filename) > 0) { gsub(".*filename=\"?([^\"]+)\"?", "\\1", filename) } else { basename(resp$url) } # 下载文件 curl_download(link, destfile = filename, handle = h)
内容的提问来源于stack exchange,提问作者ffsffs
相关产品推荐
相关产品推荐

