You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R语言中批量下载多URL指向的JPG图片?

批量下载图片URL列表中的JPG图片

感谢各位对我之前问题的帮助!我现在有一个新问题:如何使用循环函数批量下载数据框中指向JPG图片的URL链接对应的图片?我已编写好获取图片URL的代码如下:

# load libraries and packages
library("rvest")
library("ralger")
library("tidyverse")
library("jpeg")
library("here")

# set the number of pages
num_pages <- 5

# set working directory for photos to be stored
setwd("~/Desktop/lab/male_generic")

# create a list to hold the output
male <- vector("list", num_pages)

# looping the scraping, images from istockphoto
for(page_result in 1:num_pages){
  link = paste0("https://www.istockphoto.com/search/2/image?alloweduse=availableforalluses&mediatype=photography&phrase=man&page=", page_result)
  male[[page_result]] <- images_preview(link)
}

male <- unlist(male)

目前我仅能实现单张图片下载,代码如下:

test = "https://media.istockphoto.com/id/1028900652/photo/man-meditating-yoga-at-sunset-mountains-travel-lifestyle-relaxation-emotional-concept.jpg?s=612x612&w=0&k=20&c=96TlYdSI8POnOrcqH10GlPgOeWFjEIoY-7G_yMV4Eco="

download.file(test,'test.jpg', mode = 'wb')

希望学习批量下载的实现方法。


方法1:基础for循环批量下载

基于你已经获取的male URL向量,直接遍历每个链接,生成唯一文件名后下载:

# 遍历所有图片URL
for (i in seq_along(male)) {
  # 从URL提取istock图片ID作为文件名,避免重复覆盖
  img_id <- stringr::str_extract(male[i], "id/([0-9]+)/", group = 1)
  filename <- paste0(img_id, ".jpg")
  
  # 下载图片,mode='wb'必须设置,保证二进制文件下载完整
  download.file(url = male[i], destfile = filename, mode = "wb", quiet = TRUE)
  
  # 打印进度,方便查看下载情况
  cat("已完成第", i, "张:", filename, "\n")
}

方法2:tidyverse风格批量下载

如果习惯用tidyverse工具链,可以用purrr::walk来执行下载操作(适合有副作用的任务):

library(purrr)
library(stringr)

# 定义下载逻辑函数
download_img <- function(url) {
  img_id <- str_extract(url, "id/([0-9]+)/", group = 1)
  filename <- paste0(img_id, ".jpg")
  download.file(url, filename, mode = "wb", quiet = TRUE)
  cat("已下载:", filename, "\n")
}

# 批量执行下载
walk(male, download_img)

关键注意事项

  • 必须设置mode = "wb":JPG是二进制格式,缺失这个参数会导致图片损坏无法打开。
  • 确保文件名唯一:用图片ID、序号等命名,避免不同图片被覆盖。
  • 添加错误处理(可选):遇到失效URL时跳过错误,防止循环中断:
download_img <- function(url) {
  tryCatch({
    img_id <- str_extract(url, "id/([0-9]+)/", group = 1)
    filename <- paste0(img_id, ".jpg")
    download.file(url, filename, mode = "wb", quiet = TRUE)
    cat("已下载:", filename, "\n")
  }, error = function(e) {
    cat("下载失败:", url, ",错误信息:", e$message, "\n")
  })
}

内容的提问来源于stack exchange,提问作者Johnathan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 12:40:12