You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的tqdm实现多下载任务的单进度条展示?

整合UniProt多序列下载进度为单个进度条

原脚本可正常从UniProt下载FASTA序列,但每个序列对应独立进度条,处理数千个ID时会造成输出混乱。以下是修改后的脚本,将所有下载任务的进度整合为单个全局进度条,同时保留原有的错误ID记录功能:

修改后的完整脚本

'''
UniProt fasta下载器:从文本文件读取 accession ID,
用单个进度条展示所有序列的下载总进度,
并记录无法访问的序列ID
'''
import functools
import pathlib
import shutil
import requests
from tqdm.auto import tqdm

# 读取ID文件并整理URL列表
with open('errtest.txt', 'r') as infile:
    lines = [line.strip() for line in infile if line.strip()]  # 过滤空行

listfile_name = infile.name
file_name = listfile_name.split('.', 1)[0]
output_path = pathlib.Path(f"{file_name}seqs.fa").expanduser().resolve()
output_path.parent.mkdir(parents=True, exist_ok=True)

downloaded_count = 0
not_found = []
total_bytes = 0
valid_urls = []

# 预检查所有URL,获取有效序列的总字节数
print("正在预检查所有序列的可访问性...")
for line in lines:
    access_id = line
    url = f"https://rest.uniprot.org/uniprotkb/{access_id}.fasta"
    try:
        r = requests.head(url, allow_redirects=True)
        if r.status_code == 200:
            file_size = int(r.headers.get('Content-Length', 0))
            total_bytes += file_size
            valid_urls.append((access_id, url, file_size))
        else:
            not_found.append(access_id)
            print(f"{access_id} -- 未找到")
    except requests.exceptions.RequestException as e:
        not_found.append(access_id)
        print(f"{access_id} -- 请求失败: {str(e)}")

# 开始批量下载,使用单个全局进度条
print(f"开始下载 {len(valid_urls)} 个有效序列...")
with tqdm(total=total_bytes, desc="总下载进度", unit="B", unit_scale=True) as pbar:
    with open(output_path, "ab") as f:
        for access_id, url, file_size in valid_urls:
            r = requests.get(url, stream=True, allow_redirects=True)
            r.raw.read = functools.partial(r.raw.read, decode_content=True)
            
            # 分块写入文件并更新全局进度条
            for chunk in r.iter_content(chunk_size=8192):
                if chunk:
                    f.write(chunk)
                    pbar.update(len(chunk))
            
            downloaded_count += 1

# 输出结果统计
print(f"\n无法访问的序列ID列表:\n{not_found}")
print(f"成功下载 {downloaded_count} 个序列")

测试用ID文件(errtest.txt)

wrong1
D3VN13
B9W4V6
wrong2
A0A8S0XZH6
wrong3

典型输出示例

正在预检查所有序列的可访问性...
wrong1 -- 未找到
wrong2 -- 未找到
wrong3 -- 未找到
开始下载 3 个有效序列...
总下载进度: 100%|██████████| 1.48k/1.48k [00:00<00:00, 742kB/s]

无法访问的序列ID列表:
['wrong1', 'wrong2', 'wrong3']
成功下载 3 个序列

关键修改说明

  • 预检查阶段:通过requests.head()批量验证所有ID的可访问性,同时统计所有有效序列的总字节数,为全局进度条提供总进度基准
  • 全局进度条:使用单个tqdm实例跟踪所有下载的总字节数,每次写入文件块时实时更新进度
  • 优化文件写入:仅打开一次输出文件并持续写入,避免多次IO操作的开销
  • 异常处理:增加请求异常捕获逻辑,避免单个请求失败导致整个脚本中断

内容的提问来源于stack exchange,提问作者Irfan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 12:17:37