You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言easyPubMed逐行查询DataFrame列获取PMID的技术需求

Query PubMed for Each Row in a DataFrame & Add PMID List as a New Column (using easyPubMed)

Got it, let's break this down into actionable steps with a reproducible example you can adapt directly to your own data.

Step 1: Install & Load Required Packages

First, make sure you have the necessary tools installed. We'll use easyPubMed for PubMed queries and dplyr for smooth DataFrame manipulation:

install.packages(c("easyPubMed", "dplyr"))
library(easyPubMed)
library(dplyr)

Step 2: Set Up Your DataFrame

Replace this example with your actual data. Let's assume your DataFrame has a column named search_query containing the terms you want to look up:

# Example DataFrame
my_data <- data.frame(
  row_id = c(1, 2, 3),
  search_query = c("type 2 diabetes AND metformin",
                  "COVID-19 mRNA vaccine efficacy",
                  "Alzheimer's disease blood biomarkers")
)

Step 3: Write a Helper Function to Fetch PMIDs

This function will take a single search string, query PubMed, and return a vector of matching PMIDs. It also handles cases where no results are found:

fetch_pmid_list <- function(query) {
  # Get PubMed search IDs
  pubmed_results <- get_pubmed_ids(query)
  
  # Return empty vector if no matches
  if (pubmed_results$Count == 0) {
    return(character(0))
  }
  
  # Extract and return PMIDs from the results
  pubmed_data <- fetch_pubmed_data(pubmed_results)
  pmid_vector <- extract_pmid(pubmed_data)
  
  return(pmid_vector)
}

Step 4: Apply the Function to Every Row

Use rowwise() to process each row individually, then add the PMID list as a new column called PMID:

# Add PMID column to your DataFrame
my_data_with_pmids <- my_data %>%
  rowwise() %>%
  mutate(PMID = list(fetch_pmid_list(search_query))) %>%
  ungroup()

If you print the result, you'll see each row has a list of PMIDs (or an empty list if no matches were found):

print(my_data_with_pmids)

Key Tips for Smooth Execution

  • Avoid API Blocking: PubMed limits requests to ~3 per second. If you have a large DataFrame, add a small delay inside the fetch_pmid_list function:
    fetch_pmid_list <- function(query) {
      Sys.sleep(0.5) # Add 0.5-second delay between requests
      # Rest of the function code...
    }
    
  • Handle Empty Results: The function returns an empty vector when no matches exist. If you prefer NA instead, replace return(character(0)) with return(NA_character_).
  • Use PubMed Syntax: Your search_query column can use standard PubMed search syntax (MeSH terms, date filters, etc.)—easyPubMed interprets it just like the PubMed website.

内容的提问来源于stack exchange,提问作者taurian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:13:59