使用RSelenium多页面爬取并通过正则表达式筛选邮箱:代码优化咨询
Let's break down why your current code is slow and missing emails, then fix it with two better approaches—one refined RSelenium version, and a much faster rvest alternative (since this site doesn't require JavaScript rendering, we can skip the heavy browser simulation!).
Issues with Your Original Code
- Repeated page loads: You navigate back to the main professors page every loop iteration—this is the biggest speed killer, as you're reloading the same page dozens/hundreds of times unnecessarily.
- Unreliable email extraction: Using
str_split+grepand pickinga[2]is fragile. Page line breaks can vary, so you might be grabbing the wrong line entirely (hence missing emails). - Overhead of RSelenium: Simulating a full browser is slow compared to directly parsing HTML.
Option 1: Refined RSelenium Version (Faster, More Reliable)
We'll load the main page once, grab all professor links and names first, then loop through those links without reloading the main page. We'll also use proper element targeting for emails instead of string splitting.
library(RSelenium) library(stringr) # Initialize driver (only once!) rD <- rsDriver(browser = "firefox", port = 4545L, verbose = F) remDr <- rD[["client"]] # Load main page ONCE remDr$navigate("https://ki.se/en/research/professors-at-ki") # Grab all professor elements and their links/names professor_elems <- remDr$findElements(using = "xpath", "//strong/a") professor_names <- sapply(professor_elems, function(x) x$getElementText()[[1]]) professor_links <- sapply(professor_elems, function(x) x$getElementAttribute("href")[[1]]) # Initialize database database <- data.frame( Name = character(length(professor_names)), Email = character(length(professor_names)), Institution = rep("Karolinska Institute", length(professor_names)), stringsAsFactors = FALSE ) # Loop through each professor link for(i in seq_along(professor_names)){ tryCatch({ remDr$navigate(professor_links[i]) # Find email element (target <a> tags with @ in the text) email_elem <- remDr$findElements(using = "xpath", "//a[contains(text(), '@')]") if(length(email_elem) > 0){ email <- email_elem[[1]]$getElementText()[[1]] database$Name[i] <- professor_names[i] database$Email[i] <- str_trim(email) } else { database$Name[i] <- professor_names[i] database$Email[i] <- "No email found" } # Add a small delay to avoid overwhelming the server (optional but polite) Sys.sleep(0.5) }, error = function(e){ message(paste("Error processing", professor_names[i], ":", e$message)) database$Name[i] <- professor_names[i] database$Email[i] <- "Error accessing page" }) } # Clean up remDr$close() rD$server$stop() # View results head(database)
Option 2: rvest Alternative (Way Faster, No Browser Simulation)
This site's content is static (no JavaScript needed to load professor links or emails), so we can use rvest to parse HTML directly—this will be 10-100x faster than RSelenium.
library(rvest) library(dplyr) library(stringr) # Load main page and extract professor data main_page <- read_html("https://ki.se/en/research/professors-at-ki") professor_data <- main_page %>% html_elements("strong a") %>% tibble( Name = html_text(.), Link = html_attr(., "href") ) %>% mutate(Institution = "Karolinska Institute") # Function to extract email from a professor's page get_professor_email <- function(link){ tryCatch({ page <- read_html(link) # Target email links (look for <a> with mailto: prefix) email <- page %>% html_elements("a[href^='mailto:']") %>% html_text() %>% str_trim() %>% first() if(is_empty(email)) "No email found" else email }, error = function(e){ "Error accessing page" }) } # Add emails to the dataset professor_data <- professor_data %>% mutate(Email = sapply(Link, get_professor_email)) # View results head(professor_data)
Key Improvements in Both Approaches
- Single main page load: We grab all links/names upfront instead of reloading the page every time.
- Reliable email targeting: We use HTML element selectors (like
a[href^='mailto:']for mailto links) instead of fragile string splitting. - Error handling:
tryCatchensures the loop doesn't break if a page fails to load. - Polite scraping: Adding a small delay (in RSelenium) avoids hitting the server too hard.
内容的提问来源于stack exchange,提问作者Giulia

