如何使用pubmed_lookup批量读取URL文件爬取信息?
Hey there! Let's break down how to batch process those PubMed URLs from your .txt file using pubmed-lookup. I'll align this with the tool's official guidelines to make sure it's compliant and effective:
Step 1: Prep Your Environment & File
First, make sure you've got pubmed-lookup installed. If not, run this command in your terminal:
pip install pubmed-lookup
Next, double-check your .txt file:
- Each PubMed URL should be on its own line (empty lines are okay, we'll handle those in the script)
- URLs should follow the standard format, e.g.,
https://pubmed.ncbi.nlm.nih.gov/123456/
Step 2: Python Script for Batch Processing
Here's a complete, reusable script that reads your URL file, processes each entry, and saves the results. I've added comments to explain each part clearly:
from pubmed_lookup import PubMedLookup, Publication import json import time # Replace this with your actual email (required by NCBI to identify API requests) YOUR_EMAIL = "your.real.email@example.com" def load_pubmed_urls(file_path): """Load URLs from the text file, skipping empty lines.""" with open(file_path, 'r', encoding='utf-8') as file: # Strip whitespace and filter out empty lines urls = [line.strip() for line in file if line.strip()] return urls def batch_fetch_pubmed_data(urls): """Process each URL and extract key publication details.""" processed_data = [] for idx, url in enumerate(urls, 1): try: # Create a lookup instance for the target URL lookup = PubMedLookup(url, YOUR_EMAIL) # Fetch and parse the publication data publication = Publication(lookup) # Extract the fields you care about (customize this as needed!) pub_details = { "url": url, "title": publication.title, "authors": publication.authors, "abstract": publication.abstract, "journal": publication.journal, "publication_year": publication.year, "pubmed_id": publication.pubmed_id, "doi": publication.doi } processed_data.append(pub_details) print(f"Processed {idx}/{len(urls)}: {publication.title[:50]}...") # Add a small delay to avoid hitting NCBI's rate limits time.sleep(1) except Exception as e: print(f"Failed to process URL {idx} ({url}): {str(e)}") continue return processed_data if __name__ == "__main__": # Replace this with the path to your .txt file URL_FILE_PATH = "pubmed_urls.txt" # Load URLs and run the batch process pubmed_urls = load_pubmed_urls(URL_FILE_PATH) results = batch_fetch_pubmed_data(pubmed_urls) # Save results to a JSON file for easy access later with open("pubmed_results.json", 'w', encoding='utf-8') as outfile: json.dump(results, outfile, indent=2, ensure_ascii=False) print(f"\nDone! Successfully processed {len(results)} out of {len(pubmed_urls)} URLs. Results saved to pubmed_results.json")
Key Notes to Keep in Mind
- Critical: Don't skip providing a valid email. NCBI uses this to track API usage and will block requests that don't include one.
- Rate limiting: The
time.sleep(1)adds a 1-second delay between requests to stay within NCBI's guidelines. If you have hundreds of URLs, you might need to extend this slightly. - Customization: Adjust the
pub_detailsdictionary to include/exclude fields based on your needs—check the pubmed-lookup docs for all available Publication attributes. - Error handling: The script catches general exceptions and skips problematic URLs, so your batch process won't crash if one URL is invalid or unreachable.
内容的提问来源于stack exchange,提问作者LOVETW
相关产品推荐
相关产品推荐

