You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用df.apply()处理DataFrame时出现多个进度条的问题排查

问题:DataFrame.apply()时Jupyter中出现多个进度条的原因分析

问题场景

在Jupyter环境中运行脚本,对DataFrame使用apply()方法时,期望显示单个动态更新的进度条,但输出区域出现了多个重复的进度条;当DataFrame行数较少(如100行)时仅显示一个进度条,行数增多(如5000行)时进度条重复出现。尝试用sys.stdout.flush()清除输出,却导致运行时间大幅增加。

相关代码

import pandas as pd
import math

iterator_for_progressbar = 1

def progressBar(current, total, barLength = 20):
    percent = math.ceil(float(current) * 100 / total)
    arrow   = '■' * int(percent/100 * barLength)
    spaces  = '□' * (barLength - len(arrow))
    print('Calculating: %s%s %d %%' % (arrow, spaces, percent), end='\r')

def myf(row):
    global iterator_for_progressbar
    progressBar(iterator_for_progressbar, len(df), barLength = 20)   
    iterator_for_progressbar += 1

    row['1'] = 100

df = pd.DataFrame(index = range(0, 5000), columns = ['1','2','3','4','5'] )

df.apply(myf, axis=1)

原因分析

  • Jupyter输出渲染机制限制:Jupyter单元格的输出和终端不同,不完全支持\r回车覆盖逻辑。当短时间内输出大量内容(比如5000次print调用),前端会把每次print的内容都当作独立的输出条目,无法通过回车符覆盖之前的进度条行,最终导致多条进度条堆叠显示。而小数据量时输出次数少,前端能及时处理覆盖逻辑,所以只显示一条。
  • Pandas apply()的执行差异:处理大数据量时,apply()底层的循环执行逻辑和小数据量不同,可能存在内部分块处理或输出时机的变化,导致进度条的输出频率超出了Jupyter前端的覆盖处理能力。
  • sys.stdout.flush()的性能损耗:每次调用flush()都会强制将缓冲区内容同步到前端,5000次频繁刷新会产生大量IO交互,直接导致运行时间大幅增加,这是频繁强制刷新输出的必然代价。

简单改进方案

推荐使用专门适配Jupyter环境的进度条库tqdm,它能自动处理Jupyter的输出逻辑,避免重复进度条问题,同时性能更优:

import pandas as pd
from tqdm import tqdm

# 让tqdm适配pandas的apply方法
tqdm.pandas()

def myf(row):
    row['1'] = 100
    return row

df = pd.DataFrame(index=range(0, 5000), columns=['1','2','3','4','5'])

# 使用progress_apply替代apply,自动显示进度条
df.progress_apply(myf, axis=1)

内容的提问来源于stack exchange,提问作者Alexei Krasnikov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 02:13:24