RSelenium点击按钮后无法提取页面href属性的问题求助
问题解决方案:RSelenium爬取失败与API调用报错
一、RSelenium无法提取href的解决方法
问题分析
点击搜索按钮后species_links为空,大概率是以下原因:
- 固定等待时间
Sys.sleep(5)不够,页面动态加载未完成 - 元素定位器不够精准,未匹配到实际加载的链接元素
- 点击操作未真正触发搜索提交
具体修复步骤
替换固定等待为显式等待
循环检查目标元素是否出现,直到超时,比固定Sleep更可靠:# 显式等待目标链接加载,最多等待10秒 start_time <- Sys.time() while (length(remDr$findElements(using = "css selector", value = "a[href^='/species/']")) == 0) { if (difftime(Sys.time(), start_time, units = "secs") > 10) { stop("等待元素加载超时") } Sys.sleep(1) }优化元素定位器
手动审查页面后,若链接在特定容器内,调整CSS选择器,比如:species_links <- remDr$findElements(using = "css selector", value = ".search-results a[href^='/species/']")改用回车提交搜索
部分页面按钮点击可能存在交互问题,用回车提交更稳定:input_element$sendKeysToElement(list("Abies balsamea", key = "enter"))批量遍历物种列表的完整代码
封装成函数,自动处理每个物种的搜索和href提取:library(RSelenium) remDr <- remoteDriver( remoteServerAddr = "localhost", port = 4445L, browserName = "firefox" ) remDr$open() remDr$navigate("https://ser-sid.org/") webElem <- remDr$findElement(using = "class", "flex") input_element <- webElem$findChildElement(using = "css selector", value = "input[type='text']") get_species_href <- function(species_name) { input_element$clearElement() input_element$sendKeysToElement(list(species_name, key = "enter")) # 显式等待元素 start_time <- Sys.time() while (length(remDr$findElements(using = "css selector", value = "a[href^='/species/']")) == 0) { if (difftime(Sys.time(), start_time, units = "secs") > 10) { warning(paste("物种", species_name, "未找到结果")) return(NA) } Sys.sleep(1) } species_links <- remDr$findElements(using = "css selector", value = "a[href^='/species/']") href <- sapply(species_links, function(link) link$getElementAttribute("href"))[1] return(href) } # 批量处理物种列表 Species <- c("Abies balsamea", "Alchemilla glomerulans", "Antennaria dioica", "Atriplex glabriuscula", "Brachythecium salebrosum") results <- data.frame(Species = Species, Href = sapply(Species, get_species_href)) print(results) remDr$close()
二、API调用报错"No API key found in request"的解决方法
问题分析
Supabase的REST API需要同时提供Authorization和apikey两个请求头,之前的代码仅添加了Authorization,缺少apikey导致验证失败。
具体修复步骤
补充apikey请求头
使用与Authorization相同的anon key作为apikey(可从网页网络请求中获取),修改后的代码:library(httr) url <- "https://fyxheguykvewpdeysvoh.supabase.co/rest/v1/species_summary" params <- list( select = "id", # 仅获取id用于拼接物种链接 genus = "ilike.Abies%", epithet = "ilike.balsamea%" ) headers <- add_headers( `Content-Type` = "application/json", Authorization = "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImZ5eGhlZ3V5a3Zld3BkZXlzdm9oIiwicm9sZSI6ImFub24iLCJpYXQiOjE2NDc0MTY1MzQsImV4cCI6MTk2Mjk5MjUzNH0.XhJKVijhMUidqeTbH62zQ6r8cS6j22TYAKfbbRHMTZ8", apikey = "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImZ5eGhlZ3V5a3Zld3BkZXlzdm9oIiwicm9sZSI6ImFub24iLCJpYXQiOjE2NDc0MTY1MzQsImV4cCI6MTk2Mjk5MjUzNH0.XhJKVijhMUidqeTbH62zQ6r8cS6j22TYAKfbbRHMTZ8" ) response <- GET(url, query = params, headers = headers) if (http_type(response) == "application/json") { data <- content(response, "parsed") if (length(data) > 0) { href <- paste0("https://ser-sid.org/species/", data[[1]]$id) print(href) } else { print("未找到匹配物种") } } else { print("请求失败") }批量遍历物种列表的API版本
封装函数批量获取链接,效率比RSelenium更高:get_species_api <- function(species_name) { genus <- strsplit(species_name, " ")[[1]][1] epithet <- strsplit(species_name, " ")[[1]][2] params <- list( select = "id", genus = paste0("ilike.", genus, "%"), epithet = paste0("ilike.", epithet, "%") ) response <- GET(url, query = params, headers = headers) if (http_type(response) == "application/json") { data <- content(response, "parsed") if (length(data) > 0) { return(paste0("https://ser-sid.org/species/", data[[1]]$id)) } else { warning(paste("物种", species_name, "未找到结果")) return(NA) } } else { warning(paste("物种", species_name, "请求失败")) return(NA) } } # 批量处理 results_api <- data.frame(Species = Species, Href = sapply(Species, get_species_api)) print(results_api)
内容的提问来源于stack exchange,提问作者Derek Corcoran
相关产品推荐
相关产品推荐

