You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言网页爬取求助:爬取Qedu指定院校ENEM数据的问题

Fixing Your R Web Scraping Script for Qedu ENEM Data

Let's walk through fixing your script step by step—there are a few key issues preventing it from working correctly, plus missing pieces to actually extract usable data.

Core Issues in Your Current Script

  • Truncated XPath: Your XPath //div[@class = "span4 ... is incomplete, so R can't properly locate the elements you want to scrape.
  • Mismatched URL: The target page uses 112633-militar-de-salvador, but your script uses 112633-colegio-militar-de-salvador—this mismatch will lead to a 404 error or incorrect content.
  • Missing Data Extraction: You've fetched and parsed the page, but haven't added code to pull out the actual ENEM attendance data or return it from the function.
  • Unspecified Dependencies: Make sure you've installed and loaded all required packages (httr, rvest, xml2, dplyr).

Fixed & Complete Script

Here's a revised version with comments explaining each step:

# Install required packages if you haven't already
install.packages(c("httr", "rvest", "xml2", "dplyr"))

# Load packages
library(httr)
library(rvest)
library(xml2)
library(dplyr)

fetch_attendance <- function(year) {
  # Fix the URL to match the target page structure
  url <- paste0("http://www.qedu.org.br/escola/112633-militar-de-salvador/enem?edition=", year)
  
  tryCatch({
    # Fetch the page and throw an error if the request fails
    response <- GET(url)
    stop_for_status(response)
    
    # Parse HTML content with correct encoding
    page_content <- content(response, 'text', encoding = 'utf-8') %>%
      read_html()
    
    # Extract attendance data (adjust XPath if your target element differs)
    # Assuming data is in a table inside the "span4" div
    attendance_table <- page_content %>%
      html_node(xpath = '//div[@class="span4"]//table') %>%
      html_table(header = TRUE, fill = TRUE)
    
    # Optional: Clean up messy column names
    colnames(attendance_table) <- tolower(gsub(" ", "_", colnames(attendance_table)))
    
    # Return the cleaned data frame
    return(attendance_table)
    
  }, error = function(e) {
    # Handle errors like invalid years or missing pages
    message(paste("Error fetching data for", year, ":", e$message))
    return(NULL)
  })
}

# Test the function with 2009
enem_2009_data <- fetch_attendance(2009)
print(enem_2009_data)

Key Improvements Explained

  1. URL Correction: Fixed the school ID segment to ensure we request the exact page you want.
  2. Error Handling: Added tryCatch and stop_for_status to catch HTTP issues and give clear feedback.
  3. Complete XPath: Targeted a table within the span4 div (use your browser's dev tools to copy the exact XPath if this doesn't match the page's structure).
  4. Data Extraction: Converted the HTML table to a usable R data frame, plus optional column name cleaning for easier analysis.
  5. Function Structure: The function now returns scraped data (or NULL on error) instead of stopping mid-execution.

内容的提问来源于stack exchange,提问作者Gilbert Queiroz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:52:53