You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R爬虫无法抓取多页问答内容的技术求助

Fixing Pagination for Your Healthboards Scraper

Got it, let's get your scraper to pull all pages of answers for each thread! The core issue right now is your code only grabs the first page of each thread—we need to detect how many pages a thread has, generate links for every page, then scrape each one.

Key Changes Needed

  • Detect total pages per thread: Check the pagination navigation on each thread's first page to find how many pages exist.
  • Generate all page links: For threads with multiple pages, create URLs like thread-link-p2.html, thread-link-p3.html, etc.
  • Scrape each page: Modify your existing scraping functions to handle individual pages, then combine results across all pages of a thread.

Updated Full Code

library(rvest)
library(dplyr)
library(stringr)
library(purrr)
library(tidyr)
library(RCurl)
library(xlsx)

# Base URL for the site (to resolve relative links)
base_url <- "https://www.healthboards.com"

# Scrape thread metadata from the index page
url <- "https://www.healthboards.com/boards/aspergers-syndrome/index2.html"
h <- read_html(url)

threads <- h %>% html_nodes("#threadslist .alt1 div > a") %>% html_text()
thread_links <- h %>% html_nodes("#threadslist .alt1 div > a") %>% html_attr("href") %>% 
  paste0(base_url, .)  # Convert relative links to absolute
thread_starters <- h %>% html_nodes("#threadslist .alt1 div.smallfont") %>% 
  html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim()
views <- h %>% html_nodes(".alt2:nth-child(6)") %>% html_text() %>% 
  str_replace_all(",", "") %>% as.numeric()

# Helper function to get total number of pages for a thread
get_thread_page_count <- function(thread_url) {
  thread_html <- read_html(thread_url)
  # Look for pagination links (e.g., "1", "2", "Last")
  pagination_links <- thread_html %>% html_nodes(".pagination .smallfont a") %>% html_text()
  if (length(pagination_links) == 0) {
    return(1)  # No pagination = only 1 page
  }
  # Extract the last page number (skip "Last" link if present)
  last_page_num <- pagination_links %>% 
    str_extract("^\\d+$") %>% 
    na.omit() %>% 
    as.numeric() %>% 
    max(na.rm = TRUE)
  return(ifelse(is.na(last_page_num), 1, last_page_num))
}

# Helper function to generate all page URLs for a thread
generate_thread_pages <- function(thread_url, total_pages) {
  if (total_pages == 1) {
    return(thread_url)
  }
  # Replace the .html suffix with -pX.html for pages 2+
  base_thread_url <- str_replace(thread_url, "\\.html$", "")
  paste0(base_thread_url, "-p", 1:total_pages, ".html")
}

# Updated scraping functions (accept an html object instead of URL to avoid reloading)
scrape_posts_from_html <- function(html) {
  html %>% html_nodes(css = ".smallfont~ hr+ div") %>% 
    html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim()
}

scrape_dates_from_html <- function(html) {
  html %>% html_nodes(css = "table[id^='post'] td.thead:first-child") %>% 
    html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim()
}

scrape_author_ids_from_html <- function(html) {
  h_nodes <- html %>% html_nodes("div")
  id_index <- h_nodes %>% html_attr("id") %>% str_which(pattern = "postmenu")
  h_nodes[id_index] %>% html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim()
}

# Scrape all pages for each thread
scrape_thread_all_pages <- function(thread_url) {
  total_pages <- get_thread_page_count(thread_url)
  page_urls <- generate_thread_pages(thread_url, total_pages)
  
  # Scrape each page
  page_data <- map(page_urls, function(page_url) {
    html <- read_html(page_url)
    tibble(
      post_author_id = list(scrape_author_ids_from_html(html)),
      post = list(scrape_posts_from_html(html)),
      fec = list(scrape_dates_from_html(html))
    ) %>% unnest()
  })
  
  # Combine all pages into one dataframe
  bind_rows(page_data)
}

# Create master dataset
master_data <- tibble(threads, thread_starters, thread_links) %>% 
  mutate(thread_data = map(thread_links, scrape_thread_all_pages)) %>% 
  unnest(thread_data) %>% 
  select(threads, thread_starters, thread_links, post_author_id, post, fec)

# Export to Excel
write.xlsx(master_data, "C:/2.xlsx")

What's Changed?

  1. Absolute Links: We converted relative thread links to absolute using the base_url so the scraper can access pages correctly.
  2. Page Count Detection: The get_thread_page_count function checks the pagination bar to find how many pages a thread has. If there's no pagination, it defaults to 1.
  3. Page Link Generation: generate_thread_pages creates URLs for every page of the thread by modifying the original URL (e.g., example.html becomes example-p2.html).
  4. Per-Page Scraping: We updated the scraping functions to take an already-loaded HTML object instead of a URL, which saves time by avoiding reloading pages unnecessarily.
  5. Thread-Level Scraping: scrape_thread_all_pages handles scraping all pages for a single thread and combines the results into one dataframe.

Now your scraper will pull every post from every page of each thread—including the 3 extra posts on the second page of "Asperger's and talking to yourself"!

内容的提问来源于stack exchange,提问作者Ana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:37:38