Chrome可见但页面源码无的期刊摘要批量获取技术求助
Hey there! The issue you're hitting is that those abstracts are dynamically loaded when you click the "Show Abstract" button—they don't exist in the initial page source you pull with read_html(). That's why your rvest code returns empty strings. Let's walk through two solid solutions to get those abstracts in bulk:
Option 1: Use RSelenium to Simulate Browser Interactions
Since the abstracts only load after a user click, we can use RSelenium to mimic a real browser session, click each "Show Abstract" button, then extract the loaded content.
First, make sure you have the ChromeDriver installed (match it to your Chrome version) and the required packages:
# Install packages if you haven't already install.packages(c("RSelenium", "rvest", "dplyr")) library(RSelenium) library(rvest) library(dplyr) # Start a Chrome driver session driver <- rsDriver(browser = "chrome", port = 4567L) remDr <- driver[["client"]] # Navigate to the target page remDr$navigate("http://epubs.siam.org/toc/smjmap/38/1") # Find all "Show Abstract" buttons show_btns <- remDr$findElements(using = "css selector", value = ".showHideAbstract") # Click each button to load the abstract (add a small delay to let content load) for(btn in show_btns) { btn$clickElement() Sys.sleep(0.5) } # Grab the fully loaded page source and parse it with rvest page_html <- read_html(remDr$getPageSource()[[1]]) abstracts <- page_html %>% html_nodes(".abstractSection") %>% html_text(trim = TRUE) # Pair with your existing titles titles <- page_html %>% html_nodes(".art_title") %>% html_text(trim = TRUE) article_data <- tibble(Title = titles, Abstract = abstracts) # View the results print(article_data) # Clean up: close the browser and stop the driver remDr$close() driver$server$stop()
Option 2: Directly Call the Abstract API (Faster & More Efficient)
If you inspect the network traffic in your browser when clicking "Show Abstract", you'll see it sends an AJAX request to a dedicated endpoint for each abstract. We can skip the browser simulation entirely by grabbing these endpoints directly.
Here's how to do it:
install.packages(c("rvest", "dplyr", "httr", "stringr")) library(rvest) library(dplyr) library(httr) library(stringr) # First, scrape the initial page to get article DOIs base_url <- "http://epubs.siam.org/toc/smjmap/38/1" page_html <- read_html(base_url) # Extract DOIs from article links dois <- page_html %>% html_nodes(".art_title a") %>% html_attr("href") %>% str_extract("doi/([0-9.]+/[0-9.]+)") %>% str_remove("doi/") # Build URLs for each abstract's API endpoint abstract_endpoints <- paste0("http://epubs.siam.org/action/showAbstract?doi=", dois) # Batch request each abstract abstracts <- lapply(abstract_endpoints, function(endpoint) { resp <- GET(endpoint) if (http_status(resp)$category == "Success") { read_html(resp) %>% html_node(".abstractSection") %>% html_text(trim = TRUE) } else { NA_character_ # Mark failed requests as NA } }) %>% unlist() # Combine with titles for a complete dataset titles <- page_html %>% html_nodes(".art_title") %>% html_text(trim = TRUE) article_data <- tibble(Title = titles, Abstract = abstracts) print(article_data)
Which Option Should You Use?
- Option 2 is better for most cases: It's faster, doesn't require a browser driver, and avoids the overhead of simulating clicks.
- Option 1 is for complex scenarios: If the site changes its API or adds anti-scraping measures that block direct requests, RSelenium is more resilient.
内容的提问来源于stack exchange,提问作者syk

