You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何改进Ensembl次要等位基因爬取代码?处理多定位SNP等问题

优化后的Ensembl GRCh38次要等位基因数据爬取代码

针对原代码的三个优化需求,修改后的代码如下:

library(rvest)
library(purrr)
library(dplyr)
library(stringr)

# 需要查询的SNP列表
look <- c("rs3762444", "rs284262", "rs655598", "rs12089815", "rs12140153", 
          "rs788163", "rs1064213", "rs1106090", "rs7557796", "rs16825008")

# 定义数据爬取函数
biomart <- function(x) {
  url <- paste0("http://www.ensembl.org/homo_sapiens/Variation/Explore?v=", x)
  
  # 捕获HTTP错误及页面解析异常
  tryCatch({
    z <- read_html(url)
    
    # 提取所有染色体定位(处理多定位场景)
    loc_elements <- z %>% html_elements(xpath = '//div[contains(@class,"variation-location")]/a')
    locations <- loc_elements %>% html_text2() %>% 
      str_extract("chr\\d+:\\d+") %>% 
      str_remove("chr")  # 转换为"9:126646011"格式
    
    # 提取MAF数值
    maf_elements <- z %>% html_elements(xpath = '//span[.="Highest population MAF"]/following-sibling::span/b')
    maf_values <- maf_elements %>% html_text2()
    
    # 提取MAF来源信息
    source_elements <- z %>% html_elements(xpath = '//span[.="Highest population MAF"]/following-sibling::span')
    source_values <- if (length(maf_values) > 0) {
      source_elements %>% html_attr("title") %>% map_chr(~read_html(.) %>% html_text2())
    } else {
      rep(NA_character_, length(locations))
    }
    
    # 处理定位数与MAF数不匹配的情况
    if (length(locations) > length(maf_values)) {
      maf_values <- rep(maf_values[1], length(locations))
      source_values <- rep(source_values[1], length(locations))
    }
    
    # 返回单SNP的多定位数据
    data.frame(SNP = x, 
               chromosome_location = locations,
               MAF = maf_values,
               source = source_values,
               error = NA_character_,
               stringsAsFactors = FALSE)
  }, error = function(e) {
    # 出错时保留SNP标识,记录错误信息
    data.frame(SNP = x,
               chromosome_location = NA_character_,
               MAF = NA_character_,
               source = NA_character_,
               error = as.character(e),
               stringsAsFactors = FALSE)
  })
}

# 批量处理SNP,显示进度条
out <- map(look, biomart, .progress = TRUE)

# 合并所有结果
output <- bind_rows(out)

# 查看最终输出
output

优化说明

  • 多定位rs编号处理:改用html_elements提取所有定位节点,每个定位生成一行数据;当定位数量与MAF数量不一致时,默认所有定位共享同一MAF信息。
  • 新增染色体定位列:通过正则提取并格式化位置信息,输出格式统一为染色体号:位置(如9:126646011)。
  • HTTP错误处理:用tryCatch包裹页面请求逻辑,出错时直接返回包含对应SNP和错误详情的行,避免丢失SNP标识。

内容的提问来源于stack exchange,提问作者Abiologist

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 20:33:17