You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RSelenium多页面爬取并通过正则表达式筛选邮箱:代码优化咨询

Optimizing Your Professor Email Scraper for Karolinska Institute

Let's break down why your current code is slow and missing emails, then fix it with two better approaches—one refined RSelenium version, and a much faster rvest alternative (since this site doesn't require JavaScript rendering, we can skip the heavy browser simulation!).

Issues with Your Original Code

  • Repeated page loads: You navigate back to the main professors page every loop iteration—this is the biggest speed killer, as you're reloading the same page dozens/hundreds of times unnecessarily.
  • Unreliable email extraction: Using str_split + grep and picking a[2] is fragile. Page line breaks can vary, so you might be grabbing the wrong line entirely (hence missing emails).
  • Overhead of RSelenium: Simulating a full browser is slow compared to directly parsing HTML.

Option 1: Refined RSelenium Version (Faster, More Reliable)

We'll load the main page once, grab all professor links and names first, then loop through those links without reloading the main page. We'll also use proper element targeting for emails instead of string splitting.

library(RSelenium)
library(stringr)

# Initialize driver (only once!)
rD <- rsDriver(browser = "firefox", port = 4545L, verbose = F)
remDr <- rD[["client"]]

# Load main page ONCE
remDr$navigate("https://ki.se/en/research/professors-at-ki")

# Grab all professor elements and their links/names
professor_elems <- remDr$findElements(using = "xpath", "//strong/a")
professor_names <- sapply(professor_elems, function(x) x$getElementText()[[1]])
professor_links <- sapply(professor_elems, function(x) x$getElementAttribute("href")[[1]])

# Initialize database
database <- data.frame(
  Name = character(length(professor_names)),
  Email = character(length(professor_names)),
  Institution = rep("Karolinska Institute", length(professor_names)),
  stringsAsFactors = FALSE
)

# Loop through each professor link
for(i in seq_along(professor_names)){
  tryCatch({
    remDr$navigate(professor_links[i])
    
    # Find email element (target <a> tags with @ in the text)
    email_elem <- remDr$findElements(using = "xpath", "//a[contains(text(), '@')]")
    
    if(length(email_elem) > 0){
      email <- email_elem[[1]]$getElementText()[[1]]
      database$Name[i] <- professor_names[i]
      database$Email[i] <- str_trim(email)
    } else {
      database$Name[i] <- professor_names[i]
      database$Email[i] <- "No email found"
    }
    
    # Add a small delay to avoid overwhelming the server (optional but polite)
    Sys.sleep(0.5)
  }, error = function(e){
    message(paste("Error processing", professor_names[i], ":", e$message))
    database$Name[i] <- professor_names[i]
    database$Email[i] <- "Error accessing page"
  })
}

# Clean up
remDr$close()
rD$server$stop()

# View results
head(database)

Option 2: rvest Alternative (Way Faster, No Browser Simulation)

This site's content is static (no JavaScript needed to load professor links or emails), so we can use rvest to parse HTML directly—this will be 10-100x faster than RSelenium.

library(rvest)
library(dplyr)
library(stringr)

# Load main page and extract professor data
main_page <- read_html("https://ki.se/en/research/professors-at-ki")

professor_data <- main_page %>%
  html_elements("strong a") %>%
  tibble(
    Name = html_text(.),
    Link = html_attr(., "href")
  ) %>%
  mutate(Institution = "Karolinska Institute")

# Function to extract email from a professor's page
get_professor_email <- function(link){
  tryCatch({
    page <- read_html(link)
    # Target email links (look for <a> with mailto: prefix)
    email <- page %>%
      html_elements("a[href^='mailto:']") %>%
      html_text() %>%
      str_trim() %>%
      first()
    
    if(is_empty(email)) "No email found" else email
  }, error = function(e){
    "Error accessing page"
  })
}

# Add emails to the dataset
professor_data <- professor_data %>%
  mutate(Email = sapply(Link, get_professor_email))

# View results
head(professor_data)

Key Improvements in Both Approaches

  • Single main page load: We grab all links/names upfront instead of reloading the page every time.
  • Reliable email targeting: We use HTML element selectors (like a[href^='mailto:'] for mailto links) instead of fragile string splitting.
  • Error handling: tryCatch ensures the loop doesn't break if a page fails to load.
  • Polite scraping: Adding a small delay (in RSelenium) avoids hitting the server too hard.

内容的提问来源于stack exchange,提问作者Giulia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 04:07:41