You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R的rvest爬取BigFuture普林斯顿页面数据失败求助

Troubleshooting rvest Issues When Scraping CollegeBoard's Princeton Page

Let’s break down why you’re struggling to pull those elements with html_nodes() and fix it step by step.

Common Reasons for Your Problem

First, let’s diagnose the root cause:

  • Dynamic Content: CollegeBoard’s pages heavily rely on JavaScript to load content. rvest only fetches the static HTML source (what you see when hitting Ctrl+U), but elements like the college link or international student sidebar might render after the initial page load.
  • Fragile Selectors: The ID #cpProfile_ataglance_collegeGeneralUrl_anchor could be dynamically generated or nested inside a container that doesn’t exist in the static HTML.
  • Anti-Scraping Measures: The site might block requests without a proper user-agent header, making rvest’s default request look suspicious.

Step 1: Verify if Content is Static

First, check if the elements exist in the static HTML (add a user-agent to avoid being blocked):

library(rvest)
library(httr)

url <- "https://bigfuture.collegeboard.org/college-university-search/princeton-university"
page <- GET(url, add_headers("User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")) %>% 
  read_html()

# Test the selector
page %>% html_nodes("#cpProfile_ataglance_collegeGeneralUrl_anchor")

If this returns an empty list, the content is dynamically loaded—you’ll need a tool that executes JavaScript.

Step 2: Use RSelenium to Render Dynamic Content

RSelenium simulates a real browser, so it can load all JS-rendered content. Here’s how to use it:

install.packages("RSelenium")
library(RSelenium)

# Start a Chrome browser session
driver <- rsDriver(browser = "chrome", chromever = "latest")
remDr <- driver[["client"]]

# Navigate to the page and wait for JS to load
remDr$navigate(url)
Sys.sleep(3) # Adjust wait time if content takes longer to load

# Grab the college link
college_link <- remDr$findElement(using = "css selector", value = "#cpProfile_ataglance_collegeGeneralUrl_anchor")
college_link_text <- college_link$getElementText()[[1]]
college_link_url <- college_link$getElementAttribute("href")[[1]]

# Grab the international student sidebar content
# Replace with the actual selector from Chrome DevTools (right-click > Copy > Copy selector)
international_sidebar <- remDr$findElement(using = "css selector", value = ".your-sidebar-selector-here")
international_text <- international_sidebar$getElementText()[[1]]

# Clean up the browser session
remDr$close()
driver$server$stop()

Step 3: Fallback: Extract Data from Static JSON Scripts

If you prefer not to use RSelenium, check if the data is hidden in a JSON script tag in the static HTML:

library(jsonlite)

# Extract all JSON script tags
scripts <- page %>% html_nodes("script[type='application/json']") %>% html_text()

# Search for the college URL in the scripts
for (script in scripts) {
  if (grepl("collegeGeneralUrl", script)) {
    data <- fromJSON(script)
    print(data$cpProfile_ataglance_collegeGeneralUrl_anchor)
    break
  }
}

Key Tips

  • Always set a realistic User-Agent to avoid being blocked.
  • Use Chrome DevTools’ "Network" tab to check if content loads via an API—if so, you can call that API directly with httr instead of scraping the page.
  • For lighter-weight JS execution, try the V8 package to run JavaScript code directly in R.

内容的提问来源于stack exchange,提问作者Altamash Rafiq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:13:22