You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R爬取NAU带双下拉菜单的成绩数据遇URL无变化问题求助

Solution to Scrape NAU Class Distribution Data

The issue you're facing is that this site uses dynamic form submissions (POST requests) to load the grade table—so the initial static HTML doesn't contain the table data you need. Your static scraping approach works for pages where data is present on load, but here we need to simulate selecting the year and college via a POST request to retrieve the actual table.

Here's how to adjust your code to handle this, using modern R packages (httr for HTTP requests and rvest for HTML parsing, which are more intuitive than RCurl/XML):

Step 1: Install & Load Required Packages

# Install packages if you haven't already
if (!require(httr)) install.packages("httr")
if (!require(rvest)) install.packages("rvest")

library(httr)
library(rvest)

Step 2: Define Target Parameters

base_url <- "https://www7.nau.edu/pair/reports/ClassDistribution"
years <- 2015:2019
# Replace ... with all college codes you need (ACC, ACM, ..., WGS)
colleges <- c("ACC", "ACM", "ANTH", "BIO", "CHEM", "COMM", "CS", "EDUC", "ENGL", "GEOG", "HIST", "MATH", "MUS", "PHIL", "PHYS", "POLS", "PSY", "SOC", "WGS")

Step 3: Loop Through Year/College Combinations & Scrape Data

Since you mentioned you can handle the for loop, the critical parts here are sending the POST request and extracting the first table:

# Initialize a list to store all scraped tables
all_grade_data <- list()

for (year in years) {
  for (college in colleges) {
    # Simulate submitting the form with year and college
    response <- POST(
      url = base_url,
      body = list(
        year = as.character(year),
        college = college,
        submit = "View Report" # Confirm this value by inspecting the submit button in your browser
      ),
      encode = "form" # Critical: formats the body as form data
    )
    
    # Check if the request succeeded
    if (http_status(response)$category != "Success") {
      warning(paste("Failed to fetch data for", year, "-", college))
      next
    }
    
    # Parse the HTML response
    page_html <- read_html(response)
    
    # Extract the first table inside the #pp_table div (matches your original XPath)
    grade_table <- page_html %>%
      html_element(css = "#pp_table table") %>%
      html_table(header = FALSE, stringsAsFactors = FALSE)
    
    # Add metadata to track year/college
    grade_table$Year <- year
    grade_table$College <- college
    
    # Store the table in our list
    all_grade_data[[paste(year, college, sep = "_")]] <- grade_table
  }
}

# Optional: Combine all tables into a single data frame
combined_grades <- do.call(rbind, all_grade_data)

Key Notes:

  • Confirm the Submit Button Parameter: Right-click the "View Report" button in your browser, select "Inspect", and check the name and value attributes of the button element. If it’s something like <input type="submit" name="view_btn" value="View Report">, update the body in the POST request to use view_btn = "View Report" instead of submit = "View Report".
  • Why This Works: We’re mimicking the exact action your browser takes when you select a year/college and click submit—sending a POST request with those parameters to the server, which returns the HTML containing the table.
  • Extracting the Table: Using html_element("#pp_table table") ensures we get the exact table you’re targeting (the first table inside the pp_table div), just like your original XPath did.

If you prefer to stick with XML/RCurl, you can use RCurl::postForm() instead of httr::POST, but httr and rvest are more maintainable for modern web scraping tasks.

内容的提问来源于stack exchange,提问作者Behzod A

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 19:43:10