Python脚本Uniprot批量查询性能优化及session使用报错求助
Hey there! Let's tackle this performance bottleneck with your Uniprot retrieval script—4 million IDs is a massive dataset, so we need to fix both the speed issue and that .session() error you're hitting. Here's a breakdown of what to do:
First: Why Your Current Script Is So Slow
The biggest culprit here is almost certainly repeated HTTP connection overhead. By default, each individual ID query creates a new connection to Uniprot's servers—handshaking, sending the request, closing the connection, and repeating 100 times adds up fast. Reusing a single persistent session cuts out all that redundant work.
Fixing the .session() Error
It sounds like you're on the right track with sessions, but probably using the wrong syntax for the library you're working with. Let's cover two common scenarios:
If you're using the uniprot PyPI package
Some versions of this library let you pass a requests.Session directly to the retrieve() function. Here's the correct way to implement it:
import requests from uniprot import retrieve # Create a persistent session (reuses connections across requests) session = requests.Session() # Optional: Add a user-agent header to avoid being blocked by Uniprot's filters session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"}) # Pass the session to your retrieve call batch_ids = ["EDP09046", "P12345", ...] # Your chunk of IDs results = retrieve( ids=batch_ids, session=session, fields=["accession", "organism_name", "sequence"], # Exact fields you need format="tab" # Easy to parse later )
If the library doesn't support session parameters (or you're using bioservices)
For libraries like bioservices.Uniprot, you can attach a session directly to the client instance:
from bioservices import Uniprot import requests u = Uniprot() # Replace the default session with a persistent one u.session = requests.Session() u.session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"}) # Batch query (Uniprot recommends max 1000 IDs per request) batch_ids = ["EDP09046", "P12345", ...] mapping_results = u.get_id_mapping( from_db="UniProtKB_AC-ID", to_db="UniProtKB", query=batch_ids )
If you still get errors, double-check that you're importing requests.Session correctly and that your library version supports session injection. If all else fails, you can bypass the retrieve module entirely and call Uniprot's REST API directly with a session (example below).
Critical Speed Optimizations for 4 Million IDs
Even with a session, processing 4 million IDs one batch at a time needs more tweaks:
Batch Size Matters
Uniprot's API allows up to 1000 IDs per request—stick to this limit (or 500 if you hit rate limits). Sending 1000 IDs in one request instead of 1 cuts your total request count by 99.9%!Rate Limiting & Throttling
Uniprot blocks excessive requests (usually >5 requests/second). Add a small delay between batches (0.2-0.5 seconds) to avoid getting blocked:import time # After processing a batch: time.sleep(0.3)Chunk Your ID List
Don't load all 4 million IDs into memory at once. Read your input file in chunks (e.g., 1000 lines at a time) and process each chunk separately. This prevents memory overload.Cache Duplicate IDs
If your input has duplicate IDs, store results in a dictionary or database so you don't re-request the same data twice.Direct REST API Call (If All Else Fails)
If the retrieve module is still causing issues, call Uniprot's stream API directly with a session. This gives you full control:import requests import time def fetch_uniprot_batch(ids, session): url = "https://rest.uniprot.org/uniprotkb/stream" params = { "format": "tab", "query": " OR ".join([f"accession:{id}" for id in ids]), "fields": "accession,organism_name,sequence" } response = session.get(url, params=params) response.raise_for_status() # Catch HTTP errors return response.text # Use the session in a context manager (auto-closes when done) with requests.Session() as session: session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"}) input_file = "your_id_list.txt" batch_size = 1000 batch = [] with open(input_file, "r") as f_in, open("output_data.txt", "a") as f_out: for line in f_in: uniprot_id = line.strip() if uniprot_id: batch.append(uniprot_id) if len(batch) == batch_size: # Process batch data = fetch_uniprot_batch(batch, session) f_out.write(data + "\n") batch = [] time.sleep(0.3) # Process remaining IDs in the last batch if batch: data = fetch_uniprot_batch(batch, session) f_out.write(data + "\n")
Final Tips
- Save intermediate results: Write each batch's output to file immediately so you don't lose progress if the script crashes.
- Monitor your requests: Keep an eye on Uniprot's API status for outages or rate limit changes.
- Consider async requests: For even faster processing, use
aiohttpwith async sessions (but be extra careful with rate limits—async can send requests too quickly!).
内容的提问来源于stack exchange,提问作者Oddish

