You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言批量爬取数千页面图片详情的技术问询

R爬虫批量抓取图片详情(名称、说明、署名)

针对你的需求,以下是优化后的爬虫脚本,解决了署名提取、批量循环和自动终止的问题:

核心改进点

  • 精准提取图片署名:通过XPath定位说明后的文本内容,解决<br>分隔的问题
  • 支持指定起始编号循环爬取
  • 自动识别不存在的页面(HTTP状态码非成功或无目标节点)并终止
  • 处理空值场景,生成包含空白列的完整数据表

完整代码

library(rvest)
library(dplyr)
library(httr)

# 单个图片页面爬取函数
scrape_photo <- function(photo_id) {
  # 构造目标URL
  url <- sprintf("http://fallschurchvfd.org/photovideo.asp?photo=%d", photo_id)
  
  # 发送请求并检查页面状态
  response <- GET(url)
  if (http_status(response)$category != "Success") {
    return(NULL) # 页面不存在时返回NULL
  }
  
  page <- read_html(response)
  
  # 提取图片名称(编号拼接)
  photo_name <- sprintf("%d.jpg", photo_id)
  
  # 提取图片说明,无内容时返回空字符串
  caption <- page %>% 
    html_nodes(".text7 i") %>% 
    html_text(trim = TRUE) %>% 
    ifelse(length(.) == 0, "", .)
  
  # 提取图片署名:定位.text7下<i>标签后的非空文本
  credit <- page %>% 
    html_nodes(xpath = "//div[@class='text7']/i/following-sibling::text()[normalize-space()]") %>% 
    html_text(trim = TRUE) %>% 
    ifelse(length(.) == 0, "", .)
  
  # 返回单条数据
  data.frame(
    photo_name = photo_name,
    caption = caption,
    credit = credit,
    stringsAsFactors = FALSE
  )
}

# --------------------------
# 配置爬取参数
# --------------------------
start_id <- 1  # 指定起始编号
current_id <- start_id
photo_data <- list()  # 存储爬取结果的列表

# 循环爬取直至页面不存在
while(TRUE) {
  cat(sprintf("爬取中:编号%d\n", current_id))
  single_result <- scrape_photo(current_id)
  
  if (is.null(single_result)) {
    cat(sprintf("编号%d页面不存在,停止爬取\n", current_id))
    break
  }
  
  photo_data[[length(photo_data) + 1]] <- single_result
  current_id <- current_id + 1
  
  # 可选:添加0.5秒延迟,避免请求过于频繁
  Sys.sleep(0.5)
}

# 合并数据并导出CSV
final_data <- bind_rows(photo_data)
write.csv(final_data, "photos_details.csv", row.names = FALSE)

代码说明

  1. 署名提取:使用XPath表达式//div[@class='text7']/i/following-sibling::text()[normalize-space()]精准定位说明标签后的署名文本,跳过空的换行内容
  2. 页面存在性检查:通过httr::GET获取HTTP状态码,非成功状态(如404)直接终止该次爬取并跳出循环
  3. 空值处理:对说明、署名这类可能缺失的字段,用ifelse返回空字符串,保证数据表结构完整
  4. 请求延迟:加入Sys.sleep(0.5)降低请求频率,避免触发网站反爬机制

内容的提问来源于stack exchange,提问作者DC Historian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 20:20:40