You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本Uniprot批量查询性能优化及session使用报错求助

Hey there! Let's tackle this performance bottleneck with your Uniprot retrieval script—4 million IDs is a massive dataset, so we need to fix both the speed issue and that .session() error you're hitting. Here's a breakdown of what to do:

First: Why Your Current Script Is So Slow

The biggest culprit here is almost certainly repeated HTTP connection overhead. By default, each individual ID query creates a new connection to Uniprot's servers—handshaking, sending the request, closing the connection, and repeating 100 times adds up fast. Reusing a single persistent session cuts out all that redundant work.

Fixing the .session() Error

It sounds like you're on the right track with sessions, but probably using the wrong syntax for the library you're working with. Let's cover two common scenarios:

If you're using the uniprot PyPI package

Some versions of this library let you pass a requests.Session directly to the retrieve() function. Here's the correct way to implement it:

import requests
from uniprot import retrieve

# Create a persistent session (reuses connections across requests)
session = requests.Session()
# Optional: Add a user-agent header to avoid being blocked by Uniprot's filters
session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"})

# Pass the session to your retrieve call
batch_ids = ["EDP09046", "P12345", ...]  # Your chunk of IDs
results = retrieve(
    ids=batch_ids,
    session=session,
    fields=["accession", "organism_name", "sequence"],  # Exact fields you need
    format="tab"  # Easy to parse later
)

If the library doesn't support session parameters (or you're using bioservices)

For libraries like bioservices.Uniprot, you can attach a session directly to the client instance:

from bioservices import Uniprot
import requests

u = Uniprot()
# Replace the default session with a persistent one
u.session = requests.Session()
u.session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"})

# Batch query (Uniprot recommends max 1000 IDs per request)
batch_ids = ["EDP09046", "P12345", ...]
mapping_results = u.get_id_mapping(
    from_db="UniProtKB_AC-ID",
    to_db="UniProtKB",
    query=batch_ids
)

If you still get errors, double-check that you're importing requests.Session correctly and that your library version supports session injection. If all else fails, you can bypass the retrieve module entirely and call Uniprot's REST API directly with a session (example below).

Critical Speed Optimizations for 4 Million IDs

Even with a session, processing 4 million IDs one batch at a time needs more tweaks:

  1. Batch Size Matters
    Uniprot's API allows up to 1000 IDs per request—stick to this limit (or 500 if you hit rate limits). Sending 1000 IDs in one request instead of 1 cuts your total request count by 99.9%!

  2. Rate Limiting & Throttling
    Uniprot blocks excessive requests (usually >5 requests/second). Add a small delay between batches (0.2-0.5 seconds) to avoid getting blocked:

    import time
    # After processing a batch:
    time.sleep(0.3)
    
  3. Chunk Your ID List
    Don't load all 4 million IDs into memory at once. Read your input file in chunks (e.g., 1000 lines at a time) and process each chunk separately. This prevents memory overload.

  4. Cache Duplicate IDs
    If your input has duplicate IDs, store results in a dictionary or database so you don't re-request the same data twice.

  5. Direct REST API Call (If All Else Fails)
    If the retrieve module is still causing issues, call Uniprot's stream API directly with a session. This gives you full control:

    import requests
    import time
    
    def fetch_uniprot_batch(ids, session):
        url = "https://rest.uniprot.org/uniprotkb/stream"
        params = {
            "format": "tab",
            "query": " OR ".join([f"accession:{id}" for id in ids]),
            "fields": "accession,organism_name,sequence"
        }
        response = session.get(url, params=params)
        response.raise_for_status()  # Catch HTTP errors
        return response.text
    
    # Use the session in a context manager (auto-closes when done)
    with requests.Session() as session:
        session.headers.update({"User-Agent": "MyUniprotScript/1.0 (contact@example.com)"})
        input_file = "your_id_list.txt"
        batch_size = 1000
        batch = []
        
        with open(input_file, "r") as f_in, open("output_data.txt", "a") as f_out:
            for line in f_in:
                uniprot_id = line.strip()
                if uniprot_id:
                    batch.append(uniprot_id)
                    if len(batch) == batch_size:
                        # Process batch
                        data = fetch_uniprot_batch(batch, session)
                        f_out.write(data + "\n")
                        batch = []
                        time.sleep(0.3)
            # Process remaining IDs in the last batch
            if batch:
                data = fetch_uniprot_batch(batch, session)
                f_out.write(data + "\n")
    

Final Tips

  • Save intermediate results: Write each batch's output to file immediately so you don't lose progress if the script crashes.
  • Monitor your requests: Keep an eye on Uniprot's API status for outages or rate limit changes.
  • Consider async requests: For even faster processing, use aiohttp with async sessions (but be extra careful with rate limits—async can send requests too quickly!).

内容的提问来源于stack exchange,提问作者Oddish

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:10:57