使用R语言httr包Post()方法抓取第二页href失败求助
Hey Lukas, let's break down why you're only getting the first page of event links and fix this up!
The Core Problem with Your Current Approach
The #page2 in your POST URL is a client-side anchor—it only tells your browser to jump to a specific section of the page, but it never gets sent to the server. That's why when you POST to that URL, the server still sends back the first page content instead of the second.
How to Fix It: Two Reliable Methods
Method 1: Target the AJAX Request Directly
Dynamic pages like this load additional content via an AJAX request when you click the pagination button. Here's how to find and replicate that request:
- Open your browser's Developer Tools (F12), switch to the Network tab.
- Click the "Page 2" button on the site. Look for a new request (usually under XHR/Fetch) that loads the second page content.
- Check the request details—you'll likely see a
pageparameter (or similar) being sent in the POST body.
Here's modified code using this approach (assuming the correct parameter is page=2):
library(rvest) library(httr) library(dplyr) # Fetch first page (your original logic works here) res_page1 <- GET( "http://www.pg-sanderau.de/aktuelles---termine/", user_agent("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.50 Safari/537.36") ) Links1 <- content(res_page1, "text") %>% read_html() %>% html_nodes(".itemtitle a") %>% html_attr("href") %>% data.frame(Links = .) # Fetch second page with the correct pagination parameter res_page2 <- POST( url = "http://www.pg-sanderau.de/aktuelles---termine/", encode = "form", user_agent("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_2) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/42.0.2311.50 Safari/537.36"), add_headers(Referer = "http://www.pg-sanderau.de/aktuelles---termine/"), body = list(page = "2") # Adjust this if your dev tools show a different parameter ) Links2 <- content(res_page2, "text") %>% read_html() %>% html_nodes(".itemtitle a") %>% html_attr("href") %>% data.frame(Links = .) # Combine both pages of links all_links <- bind_rows(Links1, Links2)
Method 2: Simulate a Browser with RSelenium
If the AJAX request is tricky to reverse-engineer, using browser automation lets you interact with the page just like a human would:
library(RSelenium) library(rvest) library(dplyr) # Start a Chrome browser (make sure ChromeDriver is installed on your system) driver <- rsDriver(browser = "chrome") remDr <- driver[["client"]] # Navigate to the target page remDr$navigate("http://www.pg-sanderau.de/aktuelles---termine/") # Grab first page links page1_source <- remDr$getPageSource()[[1]] Links1 <- page1_source %>% read_html() %>% html_nodes(".itemtitle a") %>% html_attr("href") %>% data.frame(Links = .) # Click the second page button (adjust the CSS selector if needed) second_page_btn <- remDr$findElement(using = "css selector", value = ".pagination a[href='#page2']") second_page_btn$clickElement() # Wait for dynamic content to load Sys.sleep(2) # Grab second page links page2_source <- remDr$getPageSource()[[1]] Links2 <- page2_source %>% read_html() %>% html_nodes(".itemtitle a") %>% html_attr("href") %>% data.frame(Links = .) # Clean up: close the browser remDr$close() driver[["server"]]$stop() # Combine results all_links <- bind_rows(Links1, Links2)
Key Takeaways
- Anchors like
#page2don't affect server responses—always check the actual network requests your browser makes when interacting with dynamic content. - For simple AJAX loads, replicating the POST request with the right parameters is efficient.
- For complex JS-rendered pages, browser simulation with RSelenium is a foolproof fallback.
内容的提问来源于stack exchange,提问作者Lukas

