使用Rvest爬取论坛时Date列导入错误的技术求助
问题:爬虫抓取论坛帖子日期错误
使用R的rvest包爬取HardwareZone三个不同主题的论坛帖子,要求按时间顺序整理,但所有帖子日期均显示为2021-09-21及之后,其中url_2对应的主题实际发布于2020年11月,日期抓取完全错误。
核心错误分析
代码的致命问题是重复赋值覆盖了HTML解析对象:
soup <- read_html(page1) soup <- read_html(page2) soup <- read_html(page3)
这三行执行后,soup最终仅保留了page3的HTML内容,循环中只爬取了第三个帖子的数据,另外两个帖子的内容完全没被处理,自然日期全部是第三个帖子的时间。
此外还有两个次要问题:
- 用
c()反复扩展向量效率极低,大数据量下会严重拖慢运行速度 - 末尾对
date列的重复格式化属于冗余操作
修复后的代码
# 加载所需包 library(tidyverse) library(rvest) library(lubridate) # 定义要爬取的帖子基础URL列表 thread_urls <- list( "https://forums.hardwarezone.com.sg/threads/companies-may-exit-singapore-if-they-do-not-have-access-to-the-complementary-foreign-manpower-they-need-tan-see-leng.6817819/page-", "https://forums.hardwarezone.com.sg/threads/glgt-you-can-see-that-f-b-jobs-are-really-not-on-top-of-the-minds-of-singaporeans.6404486/page-", "https://forums.hardwarezone.com.sg/threads/disappointing-hard-truth-the-singaporean-worker-is-more-expensive-than-ft-coz-of-cpf-even-if-paid-same-wages-from-mom-data.6493727/page-" ) # 初始化空列表存储所有帖子数据 all_posts <- list() # 遍历每个帖子 for (thread in thread_urls) { # 遍历帖子的前100页(可根据实际页数调整) for (i in 1:100) { # 构造当前页完整URL current_url <- paste0(thread, i) # 处理页面加载失败的情况 page <- tryCatch( read_html(current_url), error = function(e) { message(paste("无法加载页面:", current_url)) return(NULL) } ) if (is.null(page)) next # 提取当前页所有帖子节点 section <- html_nodes(page, "article.message") # 解析当前页的帖子数据 page_data <- map_dfr(section, function(j) { tibble( username = html_text(html_node(j, "a.username"), trim = TRUE), post = html_text(html_node(j, "div.bbWrapper"), trim = TRUE), date_str = html_text(html_node(j, "time.u-dt"), trim = TRUE), user_status = html_text(html_node(j, "h5.userTitle.message-userTitle"), trim = TRUE) ) %>% mutate( date = ifelse(date_str != "", as.Date(date_str, format = "%b %d, %Y"), NA_Date_) ) %>% select(-date_str) }) # 将当前页数据加入总列表 all_posts[[length(all_posts) + 1]] <- page_data } } # 合并所有数据并按日期排序 hardwarezone_posts <- bind_rows(all_posts) %>% mutate(date = as.Date(date, origin = "1970-01-01")) %>% arrange(date) # 查看前6行数据示例 head(hardwarezone_posts[, c("username", "post", "date")])
修复要点说明
- 避免变量覆盖:将三个帖子URL存入列表,逐个处理每个帖子的所有页面,确保每个帖子的HTML内容都被独立解析,不会被覆盖。
- 高效数据存储:使用
map_dfr()和tibble()替代向量扩展,大幅提升数据处理效率,同时结构更清晰易维护。 - 错误容错:加入
tryCatch()处理页面加载失败的情况,避免爬虫中途崩溃。 - 日期处理优化:在解析阶段直接转换日期格式,最后统一整理并按日期排序,满足按时间顺序整理的核心需求。
- 去除冗余操作:删除不必要的日期重复格式化步骤,简化代码逻辑。
内容的提问来源于stack exchange,提问作者nesta1990
相关产品推荐
相关产品推荐

