如何在R中运行多ID的AR查询并修改脚本批量获取结果
解决Open Targets GraphQL API多ID查询的R语言实现
由于Open Targets的GraphQL API不支持一次性传入多个ID查询,我们只能通过遍历每个ID单独发起请求,再将结果合并的方式实现批量查询。以下是具体步骤和修改后的脚本:
一、核心实现思路
- 准备待查询的ID列表(可直接定义,也可从表格/CSV文件读取)
- 复用原有的单ID查询逻辑,封装成可批量调用的函数
- 遍历ID列表,逐个发送API请求并收集结果
- 解析嵌套的返回数据,整理成结构化的数据框
二、修改后的完整脚本
# 加载所需库 library(httr) library(dplyr) library(tidyr) # 1. 定义待查询的ID列表(或从表格读取:gene_ids <- read.csv("gene_ids.csv")$ensembl_id) gene_ids <- c("ENSG00000169083", "ENSG00000130203", "ENSG00000162594") # 2. 定义GraphQL查询字符串(与原单ID查询一致) query_string <- " query target($ensemblId: String!){ target(ensemblId: $ensemblId){ id approvedSymbol biotype geneticConstraint { constraintType exp obs score oe oeLower oeUpper } tractability { id modality value } } } " # 3. 封装单ID查询函数 query_single_gene <- function(gene_id) { base_url <- "https://api.platform.opentargets.org/api/v4/graphql" variables <- list("ensemblId" = gene_id) post_body <- list(query = query_string, variables = variables) # 发送请求并处理响应 response <- POST(url = base_url, body = post_body, encode = 'json') data <- content(response)$data$target # 处理嵌套字段:将geneticConstraint和tractability转换为数据框 if (!is.null(data$geneticConstraint)) { data$geneticConstraint <- list(as.data.frame(data$geneticConstraint)) } if (!is.null(data$tractability)) { data$tractability <- list(as.data.frame(data$tractability)) } return(as.data.frame(data)) } # 4. 批量查询所有ID results_list <- lapply(gene_ids, query_single_gene) # 5. 合并结果并展开嵌套字段 final_results <- bind_rows(results_list) %>% unnest(geneticConstraint, keep_empty = TRUE) %>% unnest(tractability, keep_empty = TRUE) # 打印或保存结果 print(final_results) # write.csv(final_results, "open_targets_results.csv", row.names = FALSE)
三、关键细节说明
- ID列表读取:如果你的ID存储在表格中,只需替换
gene_ids的定义,比如从CSV读取:gene_ids <- read.csv("your_file.csv")$id_column_name - 嵌套字段处理:
geneticConstraint和tractability是嵌套的列表结构,使用unnest函数可以将其展开为表格的列,方便后续分析 - 错误处理:如果需要更健壮的代码,可以在
query_single_gene函数中添加异常捕获(比如用tryCatch),避免单个ID查询失败导致整个批量任务中断
内容的提问来源于stack exchange,提问作者Jim_13
相关产品推荐
相关产品推荐

