You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页抓取编译数据时缺失值处理及批量代码优化咨询

Hey there! Let's tackle your two web scraping challenges step by step—handling missing values and streamlining batch processing. Here's a robust, efficient solution using R's rvest package:

1. Fixing Missing Value Issues (Like Empty Email Fields)

The core problem is that when html_nodes() can't locate an element (e.g., no email link on page 3), html_text() returns an empty vector. This breaks cbind() because all columns need matching lengths.

To fix this, we'll add a safety check for each field: if the element doesn't exist, we'll return NA instead of an empty value. This keeps all columns consistent in length, so your data compilation won't throw errors.

Instead of writing repetitive code for each page, we'll restructure your workflow to:

  • Store all your DOI URLs in a single vector
  • Create a reusable function to scrape one page and return a structured data frame row
  • Use batch processing tools to run this function on every URL automatically

Complete Improved Code

library(rvest)
library(purrr) # For streamlined batch processing

# Step 1: List all your DOI URLs here
doi_urls <- c(
  "http://journals.sagepub.com/doi/full/10.3102/0013189X08317501",
  "http://journals.sagepub.com/doi/abs/10.3102/0013189X037002102",
  "http://journals.sagepub.com/doi/abs/10.3102/0013189X037002104",
  # Add more URLs in this format as needed
  "http://journals.sagepub.com/doi/full/xxxx/xxxxxxx"
)

# Step 2: Build a function to scrape a single journal page
scrape_single_page <- function(url) {
  # Read the page (add a small delay if needed to avoid overloading the server)
  page <- read_html(url)
  
  # Helper function to safely extract text (returns NA if element is missing)
  safe_extract <- function(selector, xpath = NULL) {
    element <- if (!is.null(xpath)) {
      html_nodes(page, xpath = xpath)
    } else {
      html_nodes(page, selector)
    }
    if (length(element) == 0) {
      NA_character_
    } else {
      html_text(element, trim = TRUE)
    }
  }
  
  # Extract all fields and return as a data frame row
  data.frame(
    author = safe_extract('.art_authors'),
    year = safe_extract('.year'),
    journalName = safe_extract('.journalName'),
    art_title = safe_extract('.art_title'),
    volume = safe_extract('.volume'),
    page = safe_extract('.page'),
    email = safe_extract(xpath = "//a[@class = 'email']"),
    source_url = url, # Optional: keep the URL for reference
    stringsAsFactors = FALSE
  )
}

# Step 3: Batch process all URLs and combine into one dataset
final_dataset <- map_df(doi_urls, scrape_single_page)

# View the compiled results
View(final_dataset)

Key Improvements Explained

  • Missing Value Safety: The safe_extract() function ensures every field returns a value (even NA if the element is missing), so your data frame never breaks during compilation.
  • Batch Efficiency: No more copying/pasting code for each page—just add new URLs to the doi_urls vector, and the function handles the rest.
  • Maintainable Code: Using a function makes it easy to add new fields (e.g., abstract, keywords) later by updating the data.frame() call.
  • Polite Scraping: You can add Sys.sleep(1) inside the scrape_single_page function to pause between requests, which helps avoid getting blocked by the journal's server.

内容的提问来源于stack exchange,提问作者JWH2006

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 10:00:26