Jupyter Notebook下载进度无法实时打印,执行完成后才输出
问题解决:Jupyter Notebook中下载进度无法实时输出
问题重现
在Jupyter Notebook中运行分块下载代码时,下载进度无法实时打印,需等待代码执行完成后才会一次性输出所有进度信息,即使调用了sys.stdout.flush()也无效。但测试简单循环+time.sleep()的代码时,却能正常实时输出。
下载代码如下:
def _download(url, dest_path): req = requests.get(url, stream=True) req.raise_for_status() completed = 0 size = len(req.content) for chunk in req.iter_content(chunk_size=2 ** 20): fd.write(chunk) completed += 2 ** 20 print(f"Downloading: {round(completed/size*100, 1)}%") sys.stdout.flush() def get_data(): ratings_url = ("http://www2.informatik.uni-freiburg.de/" "~cziegler/BX/BX-CSV-Dump.zip") if not os.path.exists("data"): os.makedirs("data") _download(ratings_url, "data/data.zip")
问题原因
- 提前加载全部响应内容:代码中
size = len(req.content)会强制把整个文件下载到内存中,后续的iter_content只是从内存读取分块,而非实时从网络下载,导致循环瞬间完成,进度信息被批量输出。 - 文件句柄未正确初始化:代码中
fd未定义且未使用传入的dest_path参数,存在语法错误(用户可能省略了部分代码)。 - 进度计算不准确:固定每次累加
2**20,但最后一个分块大小通常会小于设定的chunk_size,会导致进度计算超过100%。
修复后的代码
import requests import os import sys def _download(url, dest_path): req = requests.get(url, stream=True) req.raise_for_status() # 从响应头获取文件总大小 size = int(req.headers.get('content-length', 0)) completed = 0 # 正确打开目标文件,用with语句自动管理资源 with open(dest_path, 'wb') as fd: for chunk in req.iter_content(chunk_size=2**20): if chunk: # 跳过空分块 fd.write(chunk) completed += len(chunk) # 用end='\r'让进度在同一行刷新,flush确保Jupyter实时输出 print(f"Downloading: {round(completed/size*100, 1)}%", end='\r') sys.stdout.flush() print("\n下载完成") def get_data(): ratings_url = "http://www2.informatik.uni-freiburg.de/~cziegler/BX/BX-CSV-Dump.zip" if not os.path.exists("data"): os.makedirs("data") _download(ratings_url, "data/data.zip") get_data()
关键修复点
- 改用
req.headers.get('content-length')获取文件大小,避免提前加载全部内容到内存。 - 使用
with open(...)正确初始化文件句柄,确保文件被正确写入。 - 按实际分块长度累加已下载大小,保证进度计算准确。
- 使用
print(..., end='\r')让进度条在同一行刷新,搭配sys.stdout.flush()确保Jupyter实时输出。
内容的提问来源于stack exchange,提问作者jenesaisquoi
相关产品推荐
相关产品推荐

