使用rvest爬取纽约时报新冠高校病例页面时html_node获取节点失败求助
Hey there! Let's break down why your rvest script isn't picking up the school nodes from that NYT page and how to fix it.
The Core Issue: Dynamic Content
The problem here is that the college list on that page is dynamically rendered with JavaScript, not part of the static HTML that read_html() fetches directly. When you copy an XPath from your browser's inspector, you're looking at the fully rendered page (after JS has run), but read_html() only gets the raw HTML sent by the server—those school nodes don't exist there yet, which is why you get {xml_missing} <NA>.
Better Solutions (No More XPath Headaches)
Since the NYT official confirmed the data is accessible directly, we can skip scraping HTML entirely and go straight to the source. Here are two reliable approaches:
1. Fetch the Raw JSON Data (Recommended)
NYT stores the college case data in a structured JSON file that's loaded via JavaScript. You can find this by using your browser's developer tools:
- Open the page, hit F12 to open DevTools, then go to the Network tab.
- Filter for "XHR" or "Fetch" requests, then refresh the page.
- Look for requests with filenames like
schools.jsonor similar—this will contain all the school data in a clean, machine-readable format.
Once you have that JSON URL, you can load it directly into R with jsonlite:
library(jsonlite) # Replace with the actual JSON URL you found in DevTools json_url <- "https://static01.nyt.com/newsgraphics/2020/covid-college-tracker/data/schools.json" college_data <- fromJSON(json_url) # Now you can access columns like school name, case count, location directly head(college_data[, c("name", "cases", "location")])
This method is way more stable than scraping HTML—you won't have to worry about XPath breaking if the page's structure changes.
2. Use RSelenium to Render JavaScript
If you really need to scrape the rendered HTML (for example, if you want additional page elements), you can use RSelenium to launch a real browser that executes the JavaScript, then scrape the fully loaded page:
library(RSelenium) library(rvest) # Start a Chrome driver (make sure you have ChromeDriver installed first) driver <- rsDriver(browser = "chrome") remDr <- driver[["client"]] # Navigate to the page and wait for JS to load remDr$navigate("https://www.nytimes.com/interactive/2020/us/covid-college-cases-tracker.html") Sys.sleep(5) # Adjust the delay if needed to ensure full loading # Get the rendered page source pg_source <- remDr$getPageSource()[[1]] pg <- read_html(pg_source) # Now your original XPath should work xpath_first_school <- '//*[@id="school100663"]' node_first_school <- html_node(pg, xpath = xpath_first_school) # Clean up: close the browser and stop the driver remDr$close() driver$server$stop()
Why Your Original Code Failed
To recap: read_html(URL) pulls the static HTML sent by the server, but the school list is generated after the page loads using JavaScript. The nodes you're targeting with XPath only exist in the browser's rendered DOM, not the raw HTML. So rvest can't see them until the JavaScript runs.
内容的提问来源于stack exchange,提问作者Nick Criswell

