You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Chrome可见但页面源码无的期刊摘要批量获取技术求助

How to Scrape Dynamically Loaded Abstracts from SIAM Journal Pages

Hey there! The issue you're hitting is that those abstracts are dynamically loaded when you click the "Show Abstract" button—they don't exist in the initial page source you pull with read_html(). That's why your rvest code returns empty strings. Let's walk through two solid solutions to get those abstracts in bulk:

Option 1: Use RSelenium to Simulate Browser Interactions

Since the abstracts only load after a user click, we can use RSelenium to mimic a real browser session, click each "Show Abstract" button, then extract the loaded content.

First, make sure you have the ChromeDriver installed (match it to your Chrome version) and the required packages:

# Install packages if you haven't already
install.packages(c("RSelenium", "rvest", "dplyr"))

library(RSelenium)
library(rvest)
library(dplyr)

# Start a Chrome driver session
driver <- rsDriver(browser = "chrome", port = 4567L)
remDr <- driver[["client"]]

# Navigate to the target page
remDr$navigate("http://epubs.siam.org/toc/smjmap/38/1")

# Find all "Show Abstract" buttons
show_btns <- remDr$findElements(using = "css selector", value = ".showHideAbstract")

# Click each button to load the abstract (add a small delay to let content load)
for(btn in show_btns) {
  btn$clickElement()
  Sys.sleep(0.5)
}

# Grab the fully loaded page source and parse it with rvest
page_html <- read_html(remDr$getPageSource()[[1]])
abstracts <- page_html %>% 
  html_nodes(".abstractSection") %>% 
  html_text(trim = TRUE)

# Pair with your existing titles
titles <- page_html %>% html_nodes(".art_title") %>% html_text(trim = TRUE)
article_data <- tibble(Title = titles, Abstract = abstracts)

# View the results
print(article_data)

# Clean up: close the browser and stop the driver
remDr$close()
driver$server$stop()

Option 2: Directly Call the Abstract API (Faster & More Efficient)

If you inspect the network traffic in your browser when clicking "Show Abstract", you'll see it sends an AJAX request to a dedicated endpoint for each abstract. We can skip the browser simulation entirely by grabbing these endpoints directly.

Here's how to do it:

install.packages(c("rvest", "dplyr", "httr", "stringr"))

library(rvest)
library(dplyr)
library(httr)
library(stringr)

# First, scrape the initial page to get article DOIs
base_url <- "http://epubs.siam.org/toc/smjmap/38/1"
page_html <- read_html(base_url)

# Extract DOIs from article links
dois <- page_html %>% 
  html_nodes(".art_title a") %>% 
  html_attr("href") %>% 
  str_extract("doi/([0-9.]+/[0-9.]+)") %>% 
  str_remove("doi/")

# Build URLs for each abstract's API endpoint
abstract_endpoints <- paste0("http://epubs.siam.org/action/showAbstract?doi=", dois)

# Batch request each abstract
abstracts <- lapply(abstract_endpoints, function(endpoint) {
  resp <- GET(endpoint)
  if (http_status(resp)$category == "Success") {
    read_html(resp) %>% 
      html_node(".abstractSection") %>% 
      html_text(trim = TRUE)
  } else {
    NA_character_ # Mark failed requests as NA
  }
}) %>% unlist()

# Combine with titles for a complete dataset
titles <- page_html %>% html_nodes(".art_title") %>% html_text(trim = TRUE)
article_data <- tibble(Title = titles, Abstract = abstracts)

print(article_data)

Which Option Should You Use?

  • Option 2 is better for most cases: It's faster, doesn't require a browser driver, and avoids the overhead of simulating clicks.
  • Option 1 is for complex scenarios: If the site changes its API or adds anti-scraping measures that block direct requests, RSelenium is more resilient.

内容的提问来源于stack exchange,提问作者syk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:14:56