You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在给定Python脚本中使用tqdm添加进度条?

用tqdm给Python脚本加实时进度条

先给你改好能直接用的代码,同时解决原脚本里的性能大坑(每次循环重复读文件),让处理速度和进度监控都到位:

修改后的完整代码

from Bio import SeqIO 
from tqdm import tqdm
import csv 
import pandas as pd 
import re 
import time

# 预加载metadata文件为字典,避免每次循环重复读文件(性能核心优化)
path_to_file = "metadata_silva_simplified.txt"
ref_to_code = {}
with open(path_to_file, "r") as file:
    for line in file:
        # 解析ref和对应的code
        pipe_first_idx = line.find("|")
        pipe_second_idx = line.find("|", pipe_first_idx + 1)
        ref = line[pipe_first_idx+1 : pipe_second_idx-1].strip()
        code = f"kraken:taxid|{line[:pipe_first_idx-1]}"
        ref_to_code[ref] = code

# 读取名称列表
df = pd.read_csv("Name.csv")
Code = []
Names = []
Missing = []

t = time.time()
# 用tqdm包裹遍历对象,自动生成进度条
for index in tqdm(df.index, total=len(df.index), desc="处理进度"):
    name = str(df["0"][index])
    # 直接从字典查询,不用再遍历文件
    if name in ref_to_code:
        Code.append(ref_to_code[name])
        Names.append(name)
    else:
        Missing.append(name)

# 保存结果
Code_dict = {'Code': Code, 'Name': Names}
dfcodes = pd.DataFrame(Code_dict) 
dfcodes.to_csv("Codes_secondpart.csv", index=False)

missing_df = pd.DataFrame(Missing)
missing_df.to_csv("Missing_secondpart.csv", index=False)

elapsed = time.time() - t
print(f"总耗时: {elapsed:.2f}秒,单样本平均耗时: {elapsed/len(df):.4f}秒")

关键修改说明

  • 进度条整合:把原有的for index in df.index改成for index in tqdm(df.index, total=len(df.index), desc="处理进度"),tqdm会自动显示当前进度、已耗时、剩余预估时间,比原来的百分比打印直观得多
  • 性能优化:原脚本每次循环都打开读取metadata_silva_simplified.txt,数据量大时IO开销极大。现在提前把文件内容解析成字典,查询时间从O(n)降到O(1),处理速度会提升几十甚至上百倍
  • 冗余代码清理:去掉了原有的counter计数、coincident标记和百分比打印逻辑,用更简洁的字典查询替代

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 02:06:01