googledrive R包下载文件遇404错误,求解决方法及替代库
环境配置
- 系统:Ubuntu Desktop 20.04
- R版本:3.6.3
- googledrive版本:1.0.1
- 目标:从包含约3000个文件的公开Google Drive文件夹(已添加到个人谷歌账户)下载匹配
storico_2012模式的CSV文件
执行代码与错误
执行以下代码尝试下载第一个匹配文件:
library(googledrive) drive_about() s2012 <- drive_ls(path = "CSV Stazione_Parametro_AnnoMese", pattern = paste0("storico_2012")) drive_download(file = as_id(s2012$id[1]))
返回404错误:
Error: Client error: (404) Not Found * domain: global * reason: notFound * message: File not found: 0B-owdnU_9_lpY0otY3FYaF9nOG8. * locationType: parameter * location: fileId Run `rlang::last_error()` to see where the error occurred. > rlang::last_error() <error/gargle_error_request_failed> Client error: (404) Not Found * message: File not found: 0B-owdnU_9_lpSGR2SGtCUnlRekk. * domain: global * reason: notFound * location: fileId * locationType: parameter Backtrace: 1. base::mapply(function(x) drive_download(file = as_id(x)), bla$id) 3. googledrive::drive_download(file = as_id(x)) 5. googledrive:::as_dribble.drive_id(file) 6. googledrive::drive_get(id = x) 8. purrr::map(as_id(id), get_one_file) 9. googledrive:::.f(.x[[i]], ...) 10. gargle::response_process(response) 11. gargle:::stop_request_failed(error_message(resp), resp) Run `rlang::last_trace()` to see the full context.
矛盾现象
drive_ls返回的结果中明确包含该文件:
s2012 # A tibble: 201 x 3 name id drive_resource * <chr> <chr> <list> 1 storico_2012_07000027_005.csv 0B-owdnU_9_lpY0otY3FYaF9nOG8 <named list [37]> 2 storico_2012_10000001_005.csv 0B-owdnU_9_lpcmFkUDYwUzh4X0k <named list [37]> 3 storico_2012_05000020_010.csv 0B-owdnU_9_lpcTlEMTFpbjJLSVE <named list [37]> 4 storico_2012_03000006_005.csv 0B-owdnU_9_lpbDJiNFZWUy1CcEU <named list [37]> 5 storico_2012_09000018_111.csv 0B-owdnU_9_lpRHlwN0JnNVNseDg <named list [37]> 6 storico_2012_07000041_005.csv 0B-owdnU_9_lpV1hINnZtSFRYaTg <named list [37]> 7 storico_2012_04000155_009.csv 0B-owdnU_9_lpMzh6a29BQ3hJbHM <named list [37]> 8 storico_2012_09000014_020.csv 0B-owdnU_9_lpS0Y0ZFIzbV9mX1U <named list [37]> 9 storico_2012_03000006_038.csv 0B-owdnU_9_lpMlpGbkpFdVdURzQ <named list [37]> 10 storico_2012_06000036_009.csv 0B-owdnU_9_lpa0kxTTBuLU83U2s <named list [37]> # … with 191 more rows
- 同文件夹内其他文件可正常下载,但
drive_ls按模式查找需多次执行才能获取全量匹配文件
解决单个文件404错误的方法
验证文件实际状态
直接在浏览器中访问https://drive.google.com/file/d/[目标文件ID]/view,确认文件是否存在、是否有权限访问。个别共享文件可能被单独设置权限,或已被删除/移动,导致drive_ls的缓存数据未更新。若文件存在但权限异常,重新将整个共享文件夹添加到个人云端硬盘(右键文件夹→"添加到我的云端硬盘"),确保子文件权限同步。跳过
drive_get直接下载drive_download默认会调用drive_get验证文件,可能因缓存或权限校验失败触发404。可直接构造下载链接,用基础R函数下载:target_file <- s2012[1, ] download_url <- paste0("https://drive.google.com/uc?export=download&id=", target_file$id) download.file(download_url, destfile = target_file$name, method = "curl", extra = "-L")其中
extra = "-L"用于处理Google Drive的重定向。重置授权缓存
执行drive_auth(reset = TRUE)重新完成谷歌账户授权,清除本地权限缓存后,重新调用drive_ls获取最新文件列表,再尝试下载。
解决drive_ls查找不全的方案
优化googledrive调用逻辑
drive_ls默认分页返回结果(每页最多1000条),单次调用无法获取全量数据。可通过循环遍历分页token获取所有匹配文件:
library(dplyr) get_full_file_list <- function(target_path, match_pattern) { # 初始化第一页结果 file_list <- drive_ls(path = target_path, pattern = match_pattern, page_size = 1000) next_token <- attr(file_list, "next_page_token") # 循环获取后续分页 while (!is.null(next_token)) { next_page <- drive_ls(path = target_path, pattern = match_pattern, page_size = 1000, page_token = next_token) file_list <- bind_rows(file_list, next_page) next_token <- attr(next_page, "next_page_token") } return(file_list) } # 获取全量2012年文件 s2012_full <- get_full_file_list("CSV Stazione_Parametro_AnnoMese", "storico_2012")
替代工具/库推荐
googlesheets4
与googledrive同属gargle生态,底层API调用逻辑更稳定,可用于Drive文件的遍历与下载,处理大文件夹分页时不易遗漏。rclone命令行工具
针对大量文件的同步场景,效率远高于R包。配置好Google Drive远程连接后,可直接通过命令批量同步指定模式的文件:# 同步所有storico_2012开头的CSV到本地文件夹 rclone sync "remote:CSV Stazione_Parametro_AnnoMese" ./2012_csv_files --include "storico_2012*.csv"在R中可通过
system()调用:system('rclone sync "remote:CSV Stazione_Parametro_AnnoMese" ./2012_csv_files --include "storico_2012*.csv"')注:需先通过
rclone config完成Google Drive的远程配置。
内容的提问来源于stack exchange,提问作者Alessandro

