如何在R中从指定ZIP链接提取以s_开头的文件?
R代码补充:提取ZIP中以
s_开头的文件 你可以基于现有代码,通过以下步骤实现仅提取目标文件,并且能轻松适配后续批量处理需求:
temp <- tempfile() download.file("https://genelab-data.ndc.nasa.gov/geode-py/ws/studies/GLDS-87/download?source=datamanager&file=GLDS-87_metadata_Zanello_STS135-ISA.zip", temp) # 1. 获取ZIP包内所有文件的列表 zip_contents <- unzip(temp, list = TRUE) # 2. 筛选出以`s_`开头的文件名(正则表达式^s_匹配字符串开头) target_files <- zip_contents$Name[grepl("^s_", zip_contents$Name)] # 3. 提取目标文件到指定目录(这里设为当前目录下的extracted_files文件夹,可自行修改) unzip(temp, files = target_files, exdir = "./extracted_files") # 可选:删除临时ZIP文件,释放空间 unlink(temp)
关键说明:
unzip(..., list = TRUE)能读取ZIP内部结构,无需提前解压全部文件,节省资源grepl("^s_", ...)是核心筛选逻辑:^强制匹配字符串开头,确保只提取以s_起始的文件- 后续批量处理时,可把这段逻辑封装成函数复用:
extract_s_files <- function(zip_url, output_dir = "./extracted_files") { temp <- tempfile() download.file(zip_url, temp) zip_contents <- unzip(temp, list = TRUE) target_files <- zip_contents$Name[grepl("^s_", zip_contents$Name)] unzip(temp, files = target_files, exdir = output_dir) unlink(temp) message(paste("提取完成,共获取", length(target_files), "个文件")) } # 批量调用示例 # extract_s_files("第一个数据集URL") # extract_s_files("第二个数据集URL")
内容的提问来源于stack exchange,提问作者Nmgh
相关产品推荐
相关产品推荐

