使用df.apply()处理DataFrame时出现多个进度条的问题排查
问题:DataFrame.apply()时Jupyter中出现多个进度条的原因分析
问题场景
在Jupyter环境中运行脚本,对DataFrame使用apply()方法时,期望显示单个动态更新的进度条,但输出区域出现了多个重复的进度条;当DataFrame行数较少(如100行)时仅显示一个进度条,行数增多(如5000行)时进度条重复出现。尝试用sys.stdout.flush()清除输出,却导致运行时间大幅增加。
相关代码
import pandas as pd import math iterator_for_progressbar = 1 def progressBar(current, total, barLength = 20): percent = math.ceil(float(current) * 100 / total) arrow = '■' * int(percent/100 * barLength) spaces = '□' * (barLength - len(arrow)) print('Calculating: %s%s %d %%' % (arrow, spaces, percent), end='\r') def myf(row): global iterator_for_progressbar progressBar(iterator_for_progressbar, len(df), barLength = 20) iterator_for_progressbar += 1 row['1'] = 100 df = pd.DataFrame(index = range(0, 5000), columns = ['1','2','3','4','5'] ) df.apply(myf, axis=1)
原因分析
- Jupyter输出渲染机制限制:Jupyter单元格的输出和终端不同,不完全支持
\r回车覆盖逻辑。当短时间内输出大量内容(比如5000次print调用),前端会把每次print的内容都当作独立的输出条目,无法通过回车符覆盖之前的进度条行,最终导致多条进度条堆叠显示。而小数据量时输出次数少,前端能及时处理覆盖逻辑,所以只显示一条。 - Pandas apply()的执行差异:处理大数据量时,
apply()底层的循环执行逻辑和小数据量不同,可能存在内部分块处理或输出时机的变化,导致进度条的输出频率超出了Jupyter前端的覆盖处理能力。 - sys.stdout.flush()的性能损耗:每次调用
flush()都会强制将缓冲区内容同步到前端,5000次频繁刷新会产生大量IO交互,直接导致运行时间大幅增加,这是频繁强制刷新输出的必然代价。
简单改进方案
推荐使用专门适配Jupyter环境的进度条库tqdm,它能自动处理Jupyter的输出逻辑,避免重复进度条问题,同时性能更优:
import pandas as pd from tqdm import tqdm # 让tqdm适配pandas的apply方法 tqdm.pandas() def myf(row): row['1'] = 100 return row df = pd.DataFrame(index=range(0, 5000), columns=['1','2','3','4','5']) # 使用progress_apply替代apply,自动显示进度条 df.progress_apply(myf, axis=1)
内容的提问来源于stack exchange,提问作者Alexei Krasnikov
相关产品推荐
相关产品推荐

