R语言网页爬取求助:爬取Qedu指定院校ENEM数据的问题
Fixing Your R Web Scraping Script for Qedu ENEM Data
Let's walk through fixing your script step by step—there are a few key issues preventing it from working correctly, plus missing pieces to actually extract usable data.
Core Issues in Your Current Script
- Truncated XPath: Your XPath
//div[@class = "span4 ...is incomplete, so R can't properly locate the elements you want to scrape. - Mismatched URL: The target page uses
112633-militar-de-salvador, but your script uses112633-colegio-militar-de-salvador—this mismatch will lead to a 404 error or incorrect content. - Missing Data Extraction: You've fetched and parsed the page, but haven't added code to pull out the actual ENEM attendance data or return it from the function.
- Unspecified Dependencies: Make sure you've installed and loaded all required packages (
httr,rvest,xml2,dplyr).
Fixed & Complete Script
Here's a revised version with comments explaining each step:
# Install required packages if you haven't already install.packages(c("httr", "rvest", "xml2", "dplyr")) # Load packages library(httr) library(rvest) library(xml2) library(dplyr) fetch_attendance <- function(year) { # Fix the URL to match the target page structure url <- paste0("http://www.qedu.org.br/escola/112633-militar-de-salvador/enem?edition=", year) tryCatch({ # Fetch the page and throw an error if the request fails response <- GET(url) stop_for_status(response) # Parse HTML content with correct encoding page_content <- content(response, 'text', encoding = 'utf-8') %>% read_html() # Extract attendance data (adjust XPath if your target element differs) # Assuming data is in a table inside the "span4" div attendance_table <- page_content %>% html_node(xpath = '//div[@class="span4"]//table') %>% html_table(header = TRUE, fill = TRUE) # Optional: Clean up messy column names colnames(attendance_table) <- tolower(gsub(" ", "_", colnames(attendance_table))) # Return the cleaned data frame return(attendance_table) }, error = function(e) { # Handle errors like invalid years or missing pages message(paste("Error fetching data for", year, ":", e$message)) return(NULL) }) } # Test the function with 2009 enem_2009_data <- fetch_attendance(2009) print(enem_2009_data)
Key Improvements Explained
- URL Correction: Fixed the school ID segment to ensure we request the exact page you want.
- Error Handling: Added
tryCatchandstop_for_statusto catch HTTP issues and give clear feedback. - Complete XPath: Targeted a table within the
span4div (use your browser's dev tools to copy the exact XPath if this doesn't match the page's structure). - Data Extraction: Converted the HTML table to a usable R data frame, plus optional column name cleaning for easier analysis.
- Function Structure: The function now returns scraped data (or
NULLon error) instead of stopping mid-execution.
内容的提问来源于stack exchange,提问作者Gilbert Queiroz
相关产品推荐
相关产品推荐

