You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest爬取论坛文本并提取结构化信息的问题与解决

问题:从HardwareZone论坛提取结构化帖子数据

作为网页爬取新手,我尝试从新加坡HardwareZone论坛爬取文本数据,目前已成功获取帖子文本,但无法从中提取出包含用户名、帖子内容、发布日期、用户状态的结构化数据。

原爬取代码

library(tidyverse)
library(rvest)

# Scrape posts 
pages <- 1:32

hardwarezone_list=list()

for(i in seq_along(pages)){  
  hardwarezone_link<-paste0("https://forums.hardwarezone.com.sg/threads/glgt-you-can-see-that-f-b-jobs-are-really-not-on-top-of-the-minds-of-singaporeans.6404486/","page-",i)
  hardwarezone_page<-read_html(hardwarezone_link)  
  hardwarezone_list[[i]] <- hardwarezone_page  %>% html_nodes(".bbWrapper")  %>% html_text()
}
hardwarezone_table <- do.call(rbind,hardwarezone_list)
hardwarezone_table<- as.data.frame(hardwarezone_table)

# 输出数据示例
dput(hardwarezone_table[1:2,c(1,2)])

# 输出结果:
structure(list(V1 = c(" https://www.channelnewsasia.com/ne...bs-restaurant-association-13441340?cid=FBcna \n\"You can see that F&B jobs are really not on top of the minds of Singaporeans even when there's high unemployment,\" says a business owner.", 
"I guesss majority prefer to either send food or eat food .. not prepare the food. Haha"
), V2 = c("Recession and retrenchment only happen in EDMW ", 
"\n\t\n\t\t\n\t\t\t\n\t\t\t\ttokong said:\n\t\t\t\n\t\t\n\t\n\t\n\t\t\n\t\t\n\t\t\tno thanks, those people whose pop and mom are hawkers or have been hawkers will know. \nour parents will discourage us to become hawkers. better study hard and get a job.\nf and b jobs generate no values to your cv unless it is the end of the road for you.\nf and b pay is very jialat also. if the salary cannot feed your own family, why take the job?\nthose young punks who go into f and b either has the passion or enjoys the freedom of being not an employee\n\t\t\n\t\tClick to expand...\n\t\n\nyou will be shocked how much hawkers earn. even just those drink stall make kopi, teh kind and get soft drinks, ice from supplier and sell. don't mention bubble tea that one is considered quite artisanal.\nf&b has many positions, les amis executive chef also f&b, waitress also f&b, george quek also f&b. the value of CV is dependent on how a person wanna craft his career path, and not the industry."
)), row.names = 1:2, class = "data.frame")

理想结构化数据格式

我希望最终数据每行包含以下结构化信息:

username        post                                    date                     user status
tegridy_farm    why is that the case.                  3/10/2022               banned
Mackey          why                                   3/10/2022             Senior member
eric cartman    kyle is bad                      3/10/2022             banned

有效解决方案

以下是我找到的完美解决代码,可直接提取所需结构化数据:

hardwarezone_scraper <- function(page_number) {
    # 基础URL
    hardwarezone_link<-"https://forums.hardwarezone.com.sg/threads/glgt-you-can-see-that-f-b-jobs-are-really-not-on-top-of-the-minds-of-singaporeans.6404486/page-{page_number}"
    
    # 读取单页所有帖子的HTML内容
    messages <- read_html(glue::glue(hardwarezone_link)) %>%
    html_nodes(".message-inner")

    # 提取所需信息
    usernames <- messages %>%
    html_nodes(".message-name") %>%
        html_text()

    user_status <- messages %>% 
        html_nodes(".message-userTitle") %>%
        html_text()

    post_date <- messages %>%
        html_nodes(".listInline") %>%
        html_nodes(".u-dt") %>%
        html_text() %>%
        # 日期格式示例:"Nov 4, 2020"
        parse_date(format = "%b %d, %Y")

    post <- messages %>%
        html_nodes(".bbWrapper") %>%
        html_text()
    
    # 组合为结构化数据框并返回
    tibble(
        username = usernames,
        post = post,
        date = post_date,
        `user status` = user_status
    )
}

# 测试爬取第一页数据
hardwarezone_scraper(1)

内容的提问来源于stack exchange,提问作者nesta1990

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 02:21:04