如何在给定Python脚本中使用tqdm添加进度条?
用tqdm给Python脚本加实时进度条
先给你改好能直接用的代码,同时解决原脚本里的性能大坑(每次循环重复读文件),让处理速度和进度监控都到位:
修改后的完整代码
from Bio import SeqIO from tqdm import tqdm import csv import pandas as pd import re import time # 预加载metadata文件为字典,避免每次循环重复读文件(性能核心优化) path_to_file = "metadata_silva_simplified.txt" ref_to_code = {} with open(path_to_file, "r") as file: for line in file: # 解析ref和对应的code pipe_first_idx = line.find("|") pipe_second_idx = line.find("|", pipe_first_idx + 1) ref = line[pipe_first_idx+1 : pipe_second_idx-1].strip() code = f"kraken:taxid|{line[:pipe_first_idx-1]}" ref_to_code[ref] = code # 读取名称列表 df = pd.read_csv("Name.csv") Code = [] Names = [] Missing = [] t = time.time() # 用tqdm包裹遍历对象,自动生成进度条 for index in tqdm(df.index, total=len(df.index), desc="处理进度"): name = str(df["0"][index]) # 直接从字典查询,不用再遍历文件 if name in ref_to_code: Code.append(ref_to_code[name]) Names.append(name) else: Missing.append(name) # 保存结果 Code_dict = {'Code': Code, 'Name': Names} dfcodes = pd.DataFrame(Code_dict) dfcodes.to_csv("Codes_secondpart.csv", index=False) missing_df = pd.DataFrame(Missing) missing_df.to_csv("Missing_secondpart.csv", index=False) elapsed = time.time() - t print(f"总耗时: {elapsed:.2f}秒,单样本平均耗时: {elapsed/len(df):.4f}秒")
关键修改说明
- 进度条整合:把原有的
for index in df.index改成for index in tqdm(df.index, total=len(df.index), desc="处理进度"),tqdm会自动显示当前进度、已耗时、剩余预估时间,比原来的百分比打印直观得多 - 性能优化:原脚本每次循环都打开读取
metadata_silva_simplified.txt,数据量大时IO开销极大。现在提前把文件内容解析成字典,查询时间从O(n)降到O(1),处理速度会提升几十甚至上百倍 - 冗余代码清理:去掉了原有的
counter计数、coincident标记和百分比打印逻辑,用更简洁的字典查询替代
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

