You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用map(read_html)处理多个URL实现批量网页数据爬取?

批量爬取网页标题、内容和日期的解决方案

我来帮你搞定这个批量爬取的问题~你之前单URL的代码逻辑是对的,但批量处理时purrr::map的用法搞错了,导致报错。下面一步步给你讲清楚怎么改:

错误原因分析

你写的这段代码:

article <- test %>% mutate( art = purrr::map(read_html(Direct.Link)))

问题在于map的第一个参数应该是要迭代的向量(也就是你的Direct.Link列),第二个参数才是要应用的函数(read_html)。你直接把read_html(Direct.Link)传给map,相当于提前把整个向量塞给了read_html,而不是让map逐个处理每个URL,所以才会报"argument ".f" is missing"的错误。

正确的批量爬取代码

我们可以把单URL的爬取逻辑封装成一个函数,然后用map批量应用到Direct.Link列上,最后把结果展开成数据框的列:

第一步:加载必要的包

library(dplyr)
library(purrr)
library(rvest)

第二步:定义爬取单篇文章的函数

把你之前的单URL逻辑封装成函数,这样批量处理更清晰,还能方便处理异常:

get_article_info <- function(url) {
  # 读取网页
  page <- read_html(url)
  
  # 提取标题
  title <- page %>% html_node("title") %>% html_text() %>% trimws()
  
  # 提取正文(所有p标签内容拼接)
  content <- page %>% html_nodes("p") %>% html_text() %>% paste(collapse = " ") %>% trimws()
  
  # 提取日期(time标签内容,处理没有time标签的情况)
  time <- page %>% html_nodes("time") %>% html_text() %>% paste(collapse = " ") %>% trimws()
  # 如果没有time标签,返回NA
  if (time == "") time <- NA_character_
  
  # 返回列表,方便后续展开
  list(title = title, content = content, time = time)
}

第三步:批量处理并合并结果

用mutate+map把函数应用到每个URL,再用unnest_wider把列表列拆成单独的列:

# 你的测试数据
test <- structure(list(Participant.Name = c("Participant 1", "Participant 1", "Participant 2"), 
                       Direct.Link = c("https://chicago.suntimes.com/2020/3/26/21196297/jails-and-prisons-could-become-coronavirus-disaster", 
                                       "https://www.pressconnects.com/story/news/local/2017/07/28/cornell-study-sheds-light-students-incarcerated-parents/512160001/", 
                                       "https://www.newsobserver.com/news/local/article247133959.html" )), 
                  row.names = c(9L, 12L, 33L), class = "data.frame")

# 批量爬取并合并
test2 <- test %>%
  # 对每个URL应用爬取函数,得到包含title/content/time的列表列
  mutate(article_data = map(Direct.Link, get_article_info)) %>%
  # 把列表列展开成单独的列
  unnest_wider(article_data)

# 查看结果
print(test2)

运行这段代码后,你就能得到和你想要的test2格式完全一致的数据集了。

额外注意事项

  • 反爬机制:有些网站会阻止爬虫,你可以给read_html加个用户代理头,模拟浏览器访问:
    page <- read_html(url, user_agent = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")
    
  • 异常处理:如果某个URL打不开,函数会报错,你可以用possibly包装get_article_info,让出错时返回默认值:
    safe_get_article <- possibly(get_article_info, otherwise = list(title = NA, content = NA, time = NA))
    # 然后用map(Direct.Link, safe_get_article)
    

内容的提问来源于stack exchange,提问作者st2coker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 23:17:37