R导入可手动正常解压的合法.zip文件时报错该如何解决
问题原因
- Windows系统下
download.file()默认采用文本写入模式,下载zip等二进制文件时会自动转换换行符、插入空字节,导致压缩包结构损坏,无法被解压工具识别。 - 下载文件多出空行、fread报错嵌入空字符、无法识别csv分隔符均为压缩包损坏的典型表现。
解决方案
方案1:修改download.file的下载模式
下载时明确指定mode = "wb"(二进制写入模式)即可避免文件损坏,示例代码:
dest_path <- "你的本地存储路径" test_file <- paste0(dest_path,"/test.zip") # 明确指定二进制下载模式 download.file("https://www.portaltransparencia.gov.br/download-de-dados/despesas-execucao/202001", destfile = test_file, mode = "wb") # 解压即可正常执行 unzip(test_file, exdir = paste0(dest_path,"/unzipped")) # 读取csv dt <- data.table::fread(paste0(dest_path,"/unzipped/对应csv文件名.csv"))
方案2:使用更稳定的下载包(适配不同系统/网络环境)
如果第一种方案仍有问题,可以用httr包下载,自动处理二进制文件识别,示例代码:
library(httr) dest_path <- "你的本地存储路径" test_file <- paste0(dest_path,"/test.zip") GET("https://www.portaltransparencia.gov.br/download-de-dados/despesas-execucao/202001", write_disk(test_file, overwrite = TRUE)) unzip(test_file, exdir = paste0(dest_path,"/unzipped"))
内容的提问来源于stack exchange,提问作者Fabio Correa
相关产品推荐
相关产品推荐

