You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用pubmed_lookup批量读取URL文件爬取信息?

Hey there! Let's break down how to batch process those PubMed URLs from your .txt file using pubmed-lookup. I'll align this with the tool's official guidelines to make sure it's compliant and effective:

Step 1: Prep Your Environment & File

First, make sure you've got pubmed-lookup installed. If not, run this command in your terminal:

pip install pubmed-lookup

Next, double-check your .txt file:

  • Each PubMed URL should be on its own line (empty lines are okay, we'll handle those in the script)
  • URLs should follow the standard format, e.g., https://pubmed.ncbi.nlm.nih.gov/123456/
Step 2: Python Script for Batch Processing

Here's a complete, reusable script that reads your URL file, processes each entry, and saves the results. I've added comments to explain each part clearly:

from pubmed_lookup import PubMedLookup, Publication
import json
import time

# Replace this with your actual email (required by NCBI to identify API requests)
YOUR_EMAIL = "your.real.email@example.com"

def load_pubmed_urls(file_path):
    """Load URLs from the text file, skipping empty lines."""
    with open(file_path, 'r', encoding='utf-8') as file:
        # Strip whitespace and filter out empty lines
        urls = [line.strip() for line in file if line.strip()]
    return urls

def batch_fetch_pubmed_data(urls):
    """Process each URL and extract key publication details."""
    processed_data = []
    for idx, url in enumerate(urls, 1):
        try:
            # Create a lookup instance for the target URL
            lookup = PubMedLookup(url, YOUR_EMAIL)
            # Fetch and parse the publication data
            publication = Publication(lookup)
            
            # Extract the fields you care about (customize this as needed!)
            pub_details = {
                "url": url,
                "title": publication.title,
                "authors": publication.authors,
                "abstract": publication.abstract,
                "journal": publication.journal,
                "publication_year": publication.year,
                "pubmed_id": publication.pubmed_id,
                "doi": publication.doi
            }
            
            processed_data.append(pub_details)
            print(f"Processed {idx}/{len(urls)}: {publication.title[:50]}...")
            
            # Add a small delay to avoid hitting NCBI's rate limits
            time.sleep(1)
            
        except Exception as e:
            print(f"Failed to process URL {idx} ({url}): {str(e)}")
            continue
    
    return processed_data

if __name__ == "__main__":
    # Replace this with the path to your .txt file
    URL_FILE_PATH = "pubmed_urls.txt"
    
    # Load URLs and run the batch process
    pubmed_urls = load_pubmed_urls(URL_FILE_PATH)
    results = batch_fetch_pubmed_data(pubmed_urls)
    
    # Save results to a JSON file for easy access later
    with open("pubmed_results.json", 'w', encoding='utf-8') as outfile:
        json.dump(results, outfile, indent=2, ensure_ascii=False)
    
    print(f"\nDone! Successfully processed {len(results)} out of {len(pubmed_urls)} URLs. Results saved to pubmed_results.json")
Key Notes to Keep in Mind
  • Critical: Don't skip providing a valid email. NCBI uses this to track API usage and will block requests that don't include one.
  • Rate limiting: The time.sleep(1) adds a 1-second delay between requests to stay within NCBI's guidelines. If you have hundreds of URLs, you might need to extend this slightly.
  • Customization: Adjust the pub_details dictionary to include/exclude fields based on your needs—check the pubmed-lookup docs for all available Publication attributes.
  • Error handling: The script catches general exceptions and skips problematic URLs, so your batch process won't crash if one URL is invalid or unreachable.

内容的提问来源于stack exchange,提问作者LOVETW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 07:50:58