使用R语言rvest包批量爬取95页网页的循环实现求助
Hey Larissa, let's get that page-looping working for your document scraping task! Your single-page code is solid—we just need to wrap it in a structure that iterates through all 95 pages, collects each page's data, and combines everything into one final dataset.
Step 1: Set Up a Reusable URL Template
First, modify your base URL to include a placeholder for the page number. This lets us dynamically generate the URL for each page:
url_template <- "https://www.riksdagen.se/sv/Dokument-Lagar/?datum=2000-01-01&q=kvinn&st=2&tom=2018-12-31&doktyp=fr&p=%d"
The %d will get replaced with the page number (1 to 95) using sprintf().
Step 2: Option 1 – Use a For Loop (Straightforward)
We'll create an empty list to store each page's data frame, then loop through each page number:
library(rvest) library(dplyr) # Initialize empty list to hold page data all_pages <- list() for (page_num in 1:95) { # Generate URL for current page current_url <- sprintf(url_template, page_num) # Fetch and parse the page page <- read_html(current_url) # Extract elements (same as your original code) title_html <- html_nodes(page, '.medium-big') text_html <- html_nodes(page, '.font-bold') full_html <- html_nodes(page, '.medium-smaller') # Clean the extracted nodes text_html[[21]] <- NULL full_html[c(1, 21, 22)] <- NULL # Convert to text and create data frame title <- html_text(title_html) text <- html_text(text_html) full <- html_text(full_html) current_df <- data.frame(title, text, full) # Add current page's data to the list all_pages[[page_num]] <- current_df # Optional: Add a 1-second delay to avoid overwhelming the server Sys.sleep(1) } # Combine all page data into one final data frame final_dataset <- bind_rows(all_pages)
Step 3: Option 2 – Use Functional Programming (purrr)
If you prefer a more concise, R-native approach, use purrr::map_dfr() to loop and bind results in one step:
library(rvest) library(dplyr) library(purrr) # Create a function to scrape a single page scrape_single_page <- function(page_num) { current_url <- sprintf(url_template, page_num) page <- read_html(current_url) title_html <- html_nodes(page, '.medium-big') text_html <- html_nodes(page, '.font-bold') full_html <- html_nodes(page, '.medium-smaller') text_html[[21]] <- NULL full_html[c(1, 21, 22)] <- NULL title <- html_text(title_html) text <- html_text(text_html) full <- html_text(full_html) data.frame(title, text, full) } # Scrape all pages (with delay) and bind results final_dataset <- map_dfr(1:95, ~{ Sys.sleep(1) scrape_single_page(.x) })
Key Tips to Avoid Issues
- Add Delays: The
Sys.sleep(1)ensures you don't send too many requests too quickly, which could get your IP blocked by the site. - Handle Errors: If some pages have different HTML structures (e.g., missing elements), wrap the scraping code in
tryCatch()to skip broken pages without crashing the whole loop:scrape_single_page <- function(page_num) { tryCatch({ # ... existing scraping code ... }, error = function(e) { message(paste("Skipping page", page_num, ":", e$message)) return(data.frame(title = character(), text = character(), full = character())) }) } - Verify Selectors: Double-check that the CSS selectors (
.medium-big,.font-bold, etc.) work consistently across all pages—sometimes sites change selectors for pagination.
内容的提问来源于stack exchange,提问作者Larissa

