You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中从CSV调用URL列实现批量网页爬取的高效方案

高效批量爬取CSV中URL网页内容的R方案

嘿,我来帮你搞定这个批量爬取的需求!既然你已经把CSV文件导入R,而且里面有个叫URLs的列存了200个网页链接,完全不用费劲用c()一个个列出来——直接利用R的向量迭代功能就能高效完成批量爬取。下面给你一套实用的方案:

第一步:确认数据状态

首先确保你的数据框(假设叫df)里的URLs列是字符型,避免后续出错:

# 查看URLs列的结构
str(df$URLs)

# 如果不是字符型,转成字符型
df$URLs <- as.character(df$URLs)

第二步:准备爬取工具包

推荐用rvest(上手简单,专门用于网页爬取)和dplyr/purrr(方便数据处理),先安装加载:

# 首次使用先安装包
install.packages(c("rvest", "dplyr", "purrr"))

# 加载所需包
library(rvest)
library(dplyr)
library(purrr)

第三步:核心批量爬取代码

我们可以先定义一个爬取单页的函数,然后用迭代工具批量处理URLs列的所有链接——这样既简洁又容易维护,还能处理爬取失败的情况。

推荐方案:用purrr::map_dfr()迭代

map_dfr()会自动把每个链接的爬取结果合并成一个数据框,非常方便:

# 定义爬取单页的函数(以纽约时报为例,你需要根据目标网站调整选择器)
scrape_single_page <- function(target_url) {
  # 用tryCatch捕获错误,避免单个链接爬取失败中断整个任务
  tryCatch({
    # 读取网页内容
    page_html <- read_html(target_url)
    
    # 提取文章标题(CSS选择器可通过浏览器F12工具获取)
    article_title <- page_html %>% 
      html_element("h1") %>% 
      html_text()
    
    # 提取文章正文(同样,选择器要适配目标网站)
    article_content <- page_html %>% 
      html_elements("section[name='articleBody'] p") %>% 
      html_text() %>% 
      paste(collapse = "\n")  # 把段落合并成一段文本
    
    # 返回当前链接的爬取结果
    tibble(URL = target_url, Title = article_title, Content = article_content)
  }, error = function(e) {
    # 如果爬取失败,返回带NA的结果,方便后续排查
    tibble(URL = target_url, Title = NA_character_, Content = NA_character_)
  })
}

# 批量爬取所有URL,合并结果
final_scraped_data <- df$URLs %>% map_dfr(scrape_single_page)

# 查看前几条结果
head(final_scraped_data)

基础R替代方案:用lapply()+do.call()

如果你不想用purrr,用基础R的函数也能实现:

# 定义爬取函数
scrape_page_base <- function(target_url) {
  tryCatch({
    page_html <- read_html(target_url)
    article_title <- html_text(html_element(page_html, "h1"))
    article_content <- paste(html_text(html_elements(page_html, "section[name='articleBody'] p")), collapse = "\n")
    data.frame(URL = target_url, Title = article_title, Content = article_content, stringsAsFactors = FALSE)
  }, error = function(e) {
    data.frame(URL = target_url, Title = NA, Content = NA, stringsAsFactors = FALSE)
  })
}

# 批量爬取并合并结果
scraped_data_base <- do.call(rbind, lapply(df$URLs, scrape_page_base))

一些实用提示

  • 反爬应对:频繁请求容易被网站封禁IP,建议在爬取函数里加个延迟,比如Sys.sleep(2)(每次请求间隔2秒)
  • 选择器调整:上面的CSS选择器是针对纽约时报的,不同网站的页面结构不一样,你可以用浏览器的开发者工具(按F12)找到对应元素的选择器
  • 数据清洗:爬下来的文本可能有多余的空格或换行,可以用stringr包的str_squish()函数来清理

内容的提问来源于stack exchange,提问作者Majed Alghamdi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:54:32