R爬虫无法抓取多页问答内容的技术求助
Fixing Pagination for Your Healthboards Scraper
Got it, let's get your scraper to pull all pages of answers for each thread! The core issue right now is your code only grabs the first page of each thread—we need to detect how many pages a thread has, generate links for every page, then scrape each one.
Key Changes Needed
- Detect total pages per thread: Check the pagination navigation on each thread's first page to find how many pages exist.
- Generate all page links: For threads with multiple pages, create URLs like
thread-link-p2.html,thread-link-p3.html, etc. - Scrape each page: Modify your existing scraping functions to handle individual pages, then combine results across all pages of a thread.
Updated Full Code
library(rvest) library(dplyr) library(stringr) library(purrr) library(tidyr) library(RCurl) library(xlsx) # Base URL for the site (to resolve relative links) base_url <- "https://www.healthboards.com" # Scrape thread metadata from the index page url <- "https://www.healthboards.com/boards/aspergers-syndrome/index2.html" h <- read_html(url) threads <- h %>% html_nodes("#threadslist .alt1 div > a") %>% html_text() thread_links <- h %>% html_nodes("#threadslist .alt1 div > a") %>% html_attr("href") %>% paste0(base_url, .) # Convert relative links to absolute thread_starters <- h %>% html_nodes("#threadslist .alt1 div.smallfont") %>% html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim() views <- h %>% html_nodes(".alt2:nth-child(6)") %>% html_text() %>% str_replace_all(",", "") %>% as.numeric() # Helper function to get total number of pages for a thread get_thread_page_count <- function(thread_url) { thread_html <- read_html(thread_url) # Look for pagination links (e.g., "1", "2", "Last") pagination_links <- thread_html %>% html_nodes(".pagination .smallfont a") %>% html_text() if (length(pagination_links) == 0) { return(1) # No pagination = only 1 page } # Extract the last page number (skip "Last" link if present) last_page_num <- pagination_links %>% str_extract("^\\d+$") %>% na.omit() %>% as.numeric() %>% max(na.rm = TRUE) return(ifelse(is.na(last_page_num), 1, last_page_num)) } # Helper function to generate all page URLs for a thread generate_thread_pages <- function(thread_url, total_pages) { if (total_pages == 1) { return(thread_url) } # Replace the .html suffix with -pX.html for pages 2+ base_thread_url <- str_replace(thread_url, "\\.html$", "") paste0(base_thread_url, "-p", 1:total_pages, ".html") } # Updated scraping functions (accept an html object instead of URL to avoid reloading) scrape_posts_from_html <- function(html) { html %>% html_nodes(css = ".smallfont~ hr+ div") %>% html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim() } scrape_dates_from_html <- function(html) { html %>% html_nodes(css = "table[id^='post'] td.thead:first-child") %>% html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim() } scrape_author_ids_from_html <- function(html) { h_nodes <- html %>% html_nodes("div") id_index <- h_nodes %>% html_attr("id") %>% str_which(pattern = "postmenu") h_nodes[id_index] %>% html_text() %>% str_replace_all("\\t|\\r|\\n", "") %>% str_trim() } # Scrape all pages for each thread scrape_thread_all_pages <- function(thread_url) { total_pages <- get_thread_page_count(thread_url) page_urls <- generate_thread_pages(thread_url, total_pages) # Scrape each page page_data <- map(page_urls, function(page_url) { html <- read_html(page_url) tibble( post_author_id = list(scrape_author_ids_from_html(html)), post = list(scrape_posts_from_html(html)), fec = list(scrape_dates_from_html(html)) ) %>% unnest() }) # Combine all pages into one dataframe bind_rows(page_data) } # Create master dataset master_data <- tibble(threads, thread_starters, thread_links) %>% mutate(thread_data = map(thread_links, scrape_thread_all_pages)) %>% unnest(thread_data) %>% select(threads, thread_starters, thread_links, post_author_id, post, fec) # Export to Excel write.xlsx(master_data, "C:/2.xlsx")
What's Changed?
- Absolute Links: We converted relative thread links to absolute using the
base_urlso the scraper can access pages correctly. - Page Count Detection: The
get_thread_page_countfunction checks the pagination bar to find how many pages a thread has. If there's no pagination, it defaults to 1. - Page Link Generation:
generate_thread_pagescreates URLs for every page of the thread by modifying the original URL (e.g.,example.htmlbecomesexample-p2.html). - Per-Page Scraping: We updated the scraping functions to take an already-loaded HTML object instead of a URL, which saves time by avoiding reloading pages unnecessarily.
- Thread-Level Scraping:
scrape_thread_all_pageshandles scraping all pages for a single thread and combines the results into one dataframe.
Now your scraper will pull every post from every page of each thread—including the 3 extra posts on the second page of "Asperger's and talking to yourself"!
内容的提问来源于stack exchange,提问作者Ana
相关产品推荐
相关产品推荐

