使用R的rvest爬取BigFuture普林斯顿页面数据失败求助
Troubleshooting rvest Issues When Scraping CollegeBoard's Princeton Page
Let’s break down why you’re struggling to pull those elements with html_nodes() and fix it step by step.
Common Reasons for Your Problem
First, let’s diagnose the root cause:
- Dynamic Content: CollegeBoard’s pages heavily rely on JavaScript to load content.
rvestonly fetches the static HTML source (what you see when hitting Ctrl+U), but elements like the college link or international student sidebar might render after the initial page load. - Fragile Selectors: The ID
#cpProfile_ataglance_collegeGeneralUrl_anchorcould be dynamically generated or nested inside a container that doesn’t exist in the static HTML. - Anti-Scraping Measures: The site might block requests without a proper user-agent header, making
rvest’s default request look suspicious.
Step 1: Verify if Content is Static
First, check if the elements exist in the static HTML (add a user-agent to avoid being blocked):
library(rvest) library(httr) url <- "https://bigfuture.collegeboard.org/college-university-search/princeton-university" page <- GET(url, add_headers("User-Agent" = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")) %>% read_html() # Test the selector page %>% html_nodes("#cpProfile_ataglance_collegeGeneralUrl_anchor")
If this returns an empty list, the content is dynamically loaded—you’ll need a tool that executes JavaScript.
Step 2: Use RSelenium to Render Dynamic Content
RSelenium simulates a real browser, so it can load all JS-rendered content. Here’s how to use it:
install.packages("RSelenium") library(RSelenium) # Start a Chrome browser session driver <- rsDriver(browser = "chrome", chromever = "latest") remDr <- driver[["client"]] # Navigate to the page and wait for JS to load remDr$navigate(url) Sys.sleep(3) # Adjust wait time if content takes longer to load # Grab the college link college_link <- remDr$findElement(using = "css selector", value = "#cpProfile_ataglance_collegeGeneralUrl_anchor") college_link_text <- college_link$getElementText()[[1]] college_link_url <- college_link$getElementAttribute("href")[[1]] # Grab the international student sidebar content # Replace with the actual selector from Chrome DevTools (right-click > Copy > Copy selector) international_sidebar <- remDr$findElement(using = "css selector", value = ".your-sidebar-selector-here") international_text <- international_sidebar$getElementText()[[1]] # Clean up the browser session remDr$close() driver$server$stop()
Step 3: Fallback: Extract Data from Static JSON Scripts
If you prefer not to use RSelenium, check if the data is hidden in a JSON script tag in the static HTML:
library(jsonlite) # Extract all JSON script tags scripts <- page %>% html_nodes("script[type='application/json']") %>% html_text() # Search for the college URL in the scripts for (script in scripts) { if (grepl("collegeGeneralUrl", script)) { data <- fromJSON(script) print(data$cpProfile_ataglance_collegeGeneralUrl_anchor) break } }
Key Tips
- Always set a realistic
User-Agentto avoid being blocked. - Use Chrome DevTools’ "Network" tab to check if content loads via an API—if so, you can call that API directly with
httrinstead of scraping the page. - For lighter-weight JS execution, try the
V8package to run JavaScript code directly in R.
内容的提问来源于stack exchange,提问作者Altamash Rafiq
相关产品推荐
相关产品推荐

