You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Rvest爬取论坛时Date列导入错误的技术求助

问题:爬虫抓取论坛帖子日期错误

使用R的rvest包爬取HardwareZone三个不同主题的论坛帖子,要求按时间顺序整理,但所有帖子日期均显示为2021-09-21及之后,其中url_2对应的主题实际发布于2020年11月,日期抓取完全错误。


核心错误分析

代码的致命问题是重复赋值覆盖了HTML解析对象:

soup <- read_html(page1) 
soup <- read_html(page2) 
soup <- read_html(page3) 

这三行执行后,soup最终仅保留了page3的HTML内容,循环中只爬取了第三个帖子的数据,另外两个帖子的内容完全没被处理,自然日期全部是第三个帖子的时间。

此外还有两个次要问题:

  • 用c()反复扩展向量效率极低,大数据量下会严重拖慢运行速度
  • 末尾对date列的重复格式化属于冗余操作

修复后的代码

# 加载所需包
library(tidyverse)
library(rvest)
library(lubridate)

# 定义要爬取的帖子基础URL列表
thread_urls <- list(
  "https://forums.hardwarezone.com.sg/threads/companies-may-exit-singapore-if-they-do-not-have-access-to-the-complementary-foreign-manpower-they-need-tan-see-leng.6817819/page-",
  "https://forums.hardwarezone.com.sg/threads/glgt-you-can-see-that-f-b-jobs-are-really-not-on-top-of-the-minds-of-singaporeans.6404486/page-",
  "https://forums.hardwarezone.com.sg/threads/disappointing-hard-truth-the-singaporean-worker-is-more-expensive-than-ft-coz-of-cpf-even-if-paid-same-wages-from-mom-data.6493727/page-"
)

# 初始化空列表存储所有帖子数据
all_posts <- list()

# 遍历每个帖子
for (thread in thread_urls) {
  # 遍历帖子的前100页(可根据实际页数调整)
  for (i in 1:100) {
    # 构造当前页完整URL
    current_url <- paste0(thread, i)
    
    # 处理页面加载失败的情况
    page <- tryCatch(
      read_html(current_url),
      error = function(e) {
        message(paste("无法加载页面:", current_url))
        return(NULL)
      }
    )
    
    if (is.null(page)) next
    
    # 提取当前页所有帖子节点
    section <- html_nodes(page, "article.message")
    
    # 解析当前页的帖子数据
    page_data <- map_dfr(section, function(j) {
      tibble(
        username = html_text(html_node(j, "a.username"), trim = TRUE),
        post = html_text(html_node(j, "div.bbWrapper"), trim = TRUE),
        date_str = html_text(html_node(j, "time.u-dt"), trim = TRUE),
        user_status = html_text(html_node(j, "h5.userTitle.message-userTitle"), trim = TRUE)
      ) %>%
        mutate(
          date = ifelse(date_str != "", as.Date(date_str, format = "%b %d, %Y"), NA_Date_)
        ) %>%
        select(-date_str)
    })
    
    # 将当前页数据加入总列表
    all_posts[[length(all_posts) + 1]] <- page_data
  }
}

# 合并所有数据并按日期排序
hardwarezone_posts <- bind_rows(all_posts) %>%
  mutate(date = as.Date(date, origin = "1970-01-01")) %>%
  arrange(date)

# 查看前6行数据示例
head(hardwarezone_posts[, c("username", "post", "date")])

修复要点说明

  1. 避免变量覆盖:将三个帖子URL存入列表,逐个处理每个帖子的所有页面,确保每个帖子的HTML内容都被独立解析,不会被覆盖。
  2. 高效数据存储:使用map_dfr()和tibble()替代向量扩展,大幅提升数据处理效率,同时结构更清晰易维护。
  3. 错误容错:加入tryCatch()处理页面加载失败的情况,避免爬虫中途崩溃。
  4. 日期处理优化:在解析阶段直接转换日期格式,最后统一整理并按日期排序,满足按时间顺序整理的核心需求。
  5. 去除冗余操作:删除不必要的日期重复格式化步骤,简化代码逻辑。

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 20:38:08