You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R语言rvest包批量爬取95页网页的循环实现求助

Fixing Loop-Based Web Scraping in R (rvest)

Hey Larissa, let's get that page-looping working for your document scraping task! Your single-page code is solid—we just need to wrap it in a structure that iterates through all 95 pages, collects each page's data, and combines everything into one final dataset.

Step 1: Set Up a Reusable URL Template

First, modify your base URL to include a placeholder for the page number. This lets us dynamically generate the URL for each page:

url_template <- "https://www.riksdagen.se/sv/Dokument-Lagar/?datum=2000-01-01&q=kvinn&st=2&tom=2018-12-31&doktyp=fr&p=%d"

The %d will get replaced with the page number (1 to 95) using sprintf().

Step 2: Option 1 – Use a For Loop (Straightforward)

We'll create an empty list to store each page's data frame, then loop through each page number:

library(rvest)
library(dplyr)

# Initialize empty list to hold page data
all_pages <- list()

for (page_num in 1:95) {
  # Generate URL for current page
  current_url <- sprintf(url_template, page_num)
  
  # Fetch and parse the page
  page <- read_html(current_url)
  
  # Extract elements (same as your original code)
  title_html <- html_nodes(page, '.medium-big')
  text_html <- html_nodes(page, '.font-bold')
  full_html <- html_nodes(page, '.medium-smaller')
  
  # Clean the extracted nodes
  text_html[[21]] <- NULL
  full_html[c(1, 21, 22)] <- NULL
  
  # Convert to text and create data frame
  title <- html_text(title_html)
  text <- html_text(text_html)
  full <- html_text(full_html)
  current_df <- data.frame(title, text, full)
  
  # Add current page's data to the list
  all_pages[[page_num]] <- current_df
  
  # Optional: Add a 1-second delay to avoid overwhelming the server
  Sys.sleep(1)
}

# Combine all page data into one final data frame
final_dataset <- bind_rows(all_pages)

Step 3: Option 2 – Use Functional Programming (purrr)

If you prefer a more concise, R-native approach, use purrr::map_dfr() to loop and bind results in one step:

library(rvest)
library(dplyr)
library(purrr)

# Create a function to scrape a single page
scrape_single_page <- function(page_num) {
  current_url <- sprintf(url_template, page_num)
  page <- read_html(current_url)
  
  title_html <- html_nodes(page, '.medium-big')
  text_html <- html_nodes(page, '.font-bold')
  full_html <- html_nodes(page, '.medium-smaller')
  
  text_html[[21]] <- NULL
  full_html[c(1, 21, 22)] <- NULL
  
  title <- html_text(title_html)
  text <- html_text(text_html)
  full <- html_text(full_html)
  
  data.frame(title, text, full)
}

# Scrape all pages (with delay) and bind results
final_dataset <- map_dfr(1:95, ~{
  Sys.sleep(1)
  scrape_single_page(.x)
})

Key Tips to Avoid Issues

  • Add Delays: The Sys.sleep(1) ensures you don't send too many requests too quickly, which could get your IP blocked by the site.
  • Handle Errors: If some pages have different HTML structures (e.g., missing elements), wrap the scraping code in tryCatch() to skip broken pages without crashing the whole loop:
    scrape_single_page <- function(page_num) {
      tryCatch({
        # ... existing scraping code ...
      }, error = function(e) {
        message(paste("Skipping page", page_num, ":", e$message))
        return(data.frame(title = character(), text = character(), full = character()))
      })
    }
    
  • Verify Selectors: Double-check that the CSS selectors (.medium-big, .font-bold, etc.) work consistently across all pages—sometimes sites change selectors for pagination.

内容的提问来源于stack exchange,提问作者Larissa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 12:47:33