使用R爬取NAU带双下拉菜单的成绩数据遇URL无变化问题求助
The issue you're facing is that this site uses dynamic form submissions (POST requests) to load the grade table—so the initial static HTML doesn't contain the table data you need. Your static scraping approach works for pages where data is present on load, but here we need to simulate selecting the year and college via a POST request to retrieve the actual table.
Here's how to adjust your code to handle this, using modern R packages (httr for HTTP requests and rvest for HTML parsing, which are more intuitive than RCurl/XML):
Step 1: Install & Load Required Packages
# Install packages if you haven't already if (!require(httr)) install.packages("httr") if (!require(rvest)) install.packages("rvest") library(httr) library(rvest)
Step 2: Define Target Parameters
base_url <- "https://www7.nau.edu/pair/reports/ClassDistribution" years <- 2015:2019 # Replace ... with all college codes you need (ACC, ACM, ..., WGS) colleges <- c("ACC", "ACM", "ANTH", "BIO", "CHEM", "COMM", "CS", "EDUC", "ENGL", "GEOG", "HIST", "MATH", "MUS", "PHIL", "PHYS", "POLS", "PSY", "SOC", "WGS")
Step 3: Loop Through Year/College Combinations & Scrape Data
Since you mentioned you can handle the for loop, the critical parts here are sending the POST request and extracting the first table:
# Initialize a list to store all scraped tables all_grade_data <- list() for (year in years) { for (college in colleges) { # Simulate submitting the form with year and college response <- POST( url = base_url, body = list( year = as.character(year), college = college, submit = "View Report" # Confirm this value by inspecting the submit button in your browser ), encode = "form" # Critical: formats the body as form data ) # Check if the request succeeded if (http_status(response)$category != "Success") { warning(paste("Failed to fetch data for", year, "-", college)) next } # Parse the HTML response page_html <- read_html(response) # Extract the first table inside the #pp_table div (matches your original XPath) grade_table <- page_html %>% html_element(css = "#pp_table table") %>% html_table(header = FALSE, stringsAsFactors = FALSE) # Add metadata to track year/college grade_table$Year <- year grade_table$College <- college # Store the table in our list all_grade_data[[paste(year, college, sep = "_")]] <- grade_table } } # Optional: Combine all tables into a single data frame combined_grades <- do.call(rbind, all_grade_data)
Key Notes:
- Confirm the Submit Button Parameter: Right-click the "View Report" button in your browser, select "Inspect", and check the
nameandvalueattributes of the button element. If it’s something like<input type="submit" name="view_btn" value="View Report">, update thebodyin the POST request to useview_btn = "View Report"instead ofsubmit = "View Report". - Why This Works: We’re mimicking the exact action your browser takes when you select a year/college and click submit—sending a POST request with those parameters to the server, which returns the HTML containing the table.
- Extracting the Table: Using
html_element("#pp_table table")ensures we get the exact table you’re targeting (the first table inside thepp_tablediv), just like your original XPath did.
If you prefer to stick with XML/RCurl, you can use RCurl::postForm() instead of httr::POST, but httr and rvest are more maintainable for modern web scraping tasks.
内容的提问来源于stack exchange,提问作者Behzod A

