使用R语言批量爬取URL列表中的文章标题与内容需求
Got it, let's tackle this batch scraping task step by step. Since you already have the logic to extract title and content from a single article, scaling this to a list of URLs from a CSV is just a matter of wrapping that logic in a function and applying it across all entries. Here's how to do it:
1. Set Up Your Tools
First, make sure you have the necessary packages installed. We'll use rvest for scraping, readr to handle CSV files, dplyr for data manipulation, and purrr to safely iterate over URLs:
# Install packages if you haven't already install.packages(c("rvest", "readr", "dplyr", "purrr")) # Load them into your session library(rvest) library(readr) library(dplyr) library(purrr)
2. Read Your CSV of URLs
Assuming your CSV has a column named url (adjust the name if yours is different) that contains all the article links, read it in like this:
# Read the CSV file url_list <- read_csv("your_urls.csv") # Optional: Filter out any empty or invalid URLs to avoid errors url_list <- url_list %>% filter(!is.na(url) & url != "")
3. Create a Scraping Function
Wrap your existing single-article scraping logic into a reusable function. We'll add tryCatch to handle cases where a URL might be broken or unresponsive (so your script doesn't crash halfway through):
scrape_article <- function(article_url) { # Use tryCatch to handle errors gracefully result <- tryCatch({ # Fetch the webpage page <- read_html(article_url) # Extract title (replace with your existing selector) title <- page %>% html_element("h1") %>% html_text2() # Extract content (replace with your existing selector; adjust based on NYT's structure) content <- page %>% html_elements(".css-18sbwfn p") %>% html_text2() %>% paste(collapse = "\n") # Return a tibble with results tibble(url = article_url, title = title, content = content, status = "Success") }, error = function(e) { # If an error occurs, return NA values with an error message tibble(url = article_url, title = NA, content = NA, status = paste("Error:", e$message)) }) return(result) }
Note: Replace the CSS selectors (like "h1" or ".css-18sbwfn p") with the ones you're already using for single articles—NYT's class names might change, so double-check those if you run into issues.
4. Batch Process All URLs
Use purrr::map_dfr to apply the function to every URL in your list and combine the results into a single data frame:
# Scrape all articles (add a small delay between requests to avoid getting blocked) scraped_data <- url_list %>% mutate(scraped = map(url, ~{ Sys.sleep(2) # Wait 2 seconds between requests—adjust as needed scrape_article(.x) })) %>% unnest(scraped)
Adding a delay (Sys.sleep()) is crucial to avoid overwhelming the server and getting your IP blocked. Most websites appreciate this courtesy!
5. Save the Results
Finally, write the scraped data to a new CSV file so you can use it later:
write_csv(scraped_data, "scraped_articles.csv")
Quick Tips to Avoid Headaches
- Check for anti-scraping measures: Some sites block requests from R's default user-agent. You can set a custom one using
httr::user_agent()insideread_html:page <- read_html(article_url, user_agent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36")) - Test with a small subset first: Before running the full list, test the function on 2-3 URLs to make sure everything works as expected.
- Handle dynamic content: If some articles load content via JavaScript,
rvestalone won't work—you'll need to useRSeleniumorplaywrightto render the page first. But since you can scrape single articles, this probably isn't an issue here.
内容的提问来源于stack exchange,提问作者Majed Alghamdi

