You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

RSelenium点击按钮后无法提取页面href属性的问题求助

问题解决方案:RSelenium爬取失败与API调用报错

一、RSelenium无法提取href的解决方法

问题分析

点击搜索按钮后species_links为空,大概率是以下原因:

  • 固定等待时间Sys.sleep(5)不够,页面动态加载未完成
  • 元素定位器不够精准,未匹配到实际加载的链接元素
  • 点击操作未真正触发搜索提交

具体修复步骤

  1. 替换固定等待为显式等待
    循环检查目标元素是否出现,直到超时,比固定Sleep更可靠:

    # 显式等待目标链接加载,最多等待10秒
    start_time <- Sys.time()
    while (length(remDr$findElements(using = "css selector", value = "a[href^='/species/']")) == 0) {
      if (difftime(Sys.time(), start_time, units = "secs") > 10) {
        stop("等待元素加载超时")
      }
      Sys.sleep(1)
    }
    
  2. 优化元素定位器
    手动审查页面后,若链接在特定容器内,调整CSS选择器,比如:

    species_links <- remDr$findElements(using = "css selector", value = ".search-results a[href^='/species/']")
    
  3. 改用回车提交搜索
    部分页面按钮点击可能存在交互问题,用回车提交更稳定:

    input_element$sendKeysToElement(list("Abies balsamea", key = "enter"))
    
  4. 批量遍历物种列表的完整代码
    封装成函数,自动处理每个物种的搜索和href提取:

    library(RSelenium)
    remDr <- remoteDriver(
      remoteServerAddr = "localhost",
      port = 4445L,
      browserName = "firefox"
    )
    remDr$open()
    remDr$navigate("https://ser-sid.org/")
    
    webElem <- remDr$findElement(using = "class", "flex")
    input_element <- webElem$findChildElement(using = "css selector", value = "input[type='text']")
    
    get_species_href <- function(species_name) {
      input_element$clearElement()
      input_element$sendKeysToElement(list(species_name, key = "enter"))
      
      # 显式等待元素
      start_time <- Sys.time()
      while (length(remDr$findElements(using = "css selector", value = "a[href^='/species/']")) == 0) {
        if (difftime(Sys.time(), start_time, units = "secs") > 10) {
          warning(paste("物种", species_name, "未找到结果"))
          return(NA)
        }
        Sys.sleep(1)
      }
      
      species_links <- remDr$findElements(using = "css selector", value = "a[href^='/species/']")
      href <- sapply(species_links, function(link) link$getElementAttribute("href"))[1]
      return(href)
    }
    
    # 批量处理物种列表
    Species <- c("Abies balsamea", "Alchemilla glomerulans", "Antennaria dioica", 
                 "Atriplex glabriuscula", "Brachythecium salebrosum")
    results <- data.frame(Species = Species, Href = sapply(Species, get_species_href))
    print(results)
    
    remDr$close()
    

二、API调用报错"No API key found in request"的解决方法

问题分析

Supabase的REST API需要同时提供Authorization和apikey两个请求头,之前的代码仅添加了Authorization,缺少apikey导致验证失败。

具体修复步骤

  1. 补充apikey请求头
    使用与Authorization相同的anon key作为apikey(可从网页网络请求中获取),修改后的代码:

    library(httr)
    
    url <- "https://fyxheguykvewpdeysvoh.supabase.co/rest/v1/species_summary"
    
    params <- list(
      select = "id", # 仅获取id用于拼接物种链接
      genus = "ilike.Abies%",
      epithet = "ilike.balsamea%"
    )
    
    headers <- add_headers(
      `Content-Type` = "application/json",
      Authorization = "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImZ5eGhlZ3V5a3Zld3BkZXlzdm9oIiwicm9sZSI6ImFub24iLCJpYXQiOjE2NDc0MTY1MzQsImV4cCI6MTk2Mjk5MjUzNH0.XhJKVijhMUidqeTbH62zQ6r8cS6j22TYAKfbbRHMTZ8",
      apikey = "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJpc3MiOiJzdXBhYmFzZSIsInJlZiI6ImZ5eGhlZ3V5a3Zld3BkZXlzdm9oIiwicm9sZSI6ImFub24iLCJpYXQiOjE2NDc0MTY1MzQsImV4cCI6MTk2Mjk5MjUzNH0.XhJKVijhMUidqeTbH62zQ6r8cS6j22TYAKfbbRHMTZ8"
    )
    
    response <- GET(url, query = params, headers = headers)
    
    if (http_type(response) == "application/json") {
      data <- content(response, "parsed")
      if (length(data) > 0) {
        href <- paste0("https://ser-sid.org/species/", data[[1]]$id)
        print(href)
      } else {
        print("未找到匹配物种")
      }
    } else {
      print("请求失败")
    }
    
  2. 批量遍历物种列表的API版本
    封装函数批量获取链接,效率比RSelenium更高:

    get_species_api <- function(species_name) {
      genus <- strsplit(species_name, " ")[[1]][1]
      epithet <- strsplit(species_name, " ")[[1]][2]
      
      params <- list(
        select = "id",
        genus = paste0("ilike.", genus, "%"),
        epithet = paste0("ilike.", epithet, "%")
      )
      
      response <- GET(url, query = params, headers = headers)
      
      if (http_type(response) == "application/json") {
        data <- content(response, "parsed")
        if (length(data) > 0) {
          return(paste0("https://ser-sid.org/species/", data[[1]]$id))
        } else {
          warning(paste("物种", species_name, "未找到结果"))
          return(NA)
        }
      } else {
        warning(paste("物种", species_name, "请求失败"))
        return(NA)
      }
    }
    
    # 批量处理
    results_api <- data.frame(Species = Species, Href = sapply(Species, get_species_api))
    print(results_api)
    

内容的提问来源于stack exchange,提问作者Derek Corcoran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 03:37:34