You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow 2.6.4加载MNIST时反复输出Cleanup called...的解决方法

问题:加载MNIST数据集时持续打印"Cleanup called..."

使用TensorFlow 2.6.4加载MNIST数据集时,输出单元格反复打印"Cleanup called...",启动训练时该情况也会持续出现。

相关代码如下:

train_ds, test_ds = tfds.load(
    'mnist',
    split=['train', 'test'],
    as_supervised=True
)
def normalize_img(image, label):
    return tf.cast(image, tf.float32) / 255., label
train_ds = train_ds.map(normalize_img, num_parallel_calls=tf.data.AUTOTUNE)
train_ds = train_ds.shuffle(len(train_ds))
train_ds = train_ds.batch(128)
train_ds = train_ds.prefetch(tf.data.AUTOTUNE)
next(iter(train_ds))

输出结果:

Cleanup called...
Cleanup called...
Cleanup called...
...

已了解到相关问题的现有解决方案为降级TensorFlow,但不想采取该方式,寻求无需降级即可消除该提示的正确数据加载方法。


解决方案

以下几种调整方式可在不降级TensorFlow的前提下消除该提示:

  • 替换shuffle的buffer_size参数:避免直接使用len(train_ds)(这会触发数据集完整遍历,间接引发频繁清理),改用固定的合理buffer值:

    train_ds = train_ds.shuffle(10000)  # 替换原train_ds.shuffle(len(train_ds))
    
  • 调整并行调用数的设置:暂时将tf.data.AUTOTUNE替换为具体的固定数值,减少动态资源调度带来的清理触发:

    train_ds = train_ds.map(normalize_img, num_parallel_calls=4)
    
  • 加载数据集时禁用GCS尝试:添加try_gcs=False参数,避免与GCS相关的冗余清理逻辑触发:

    train_ds, test_ds = tfds.load(
        'mnist',
        split=['train', 'test'],
        as_supervised=True,
        try_gcs=False
    )
    
  • 添加数据集缓存:在map处理后添加缓存,减少重复的数据读取与处理操作:

    train_ds = train_ds.map(normalize_img, num_parallel_calls=tf.data.AUTOTUNE)
    train_ds = train_ds.cache()  # 新增缓存步骤
    train_ds = train_ds.shuffle(10000)
    train_ds = train_ds.batch(128)
    train_ds = train_ds.prefetch(tf.data.AUTOTUNE)
    

该提示本质是TensorFlow数据模块在资源清理时的冗余日志,并非功能错误,通过调整数据集处理流程,避免触发不必要的资源遍历或动态调度逻辑,即可彻底消除输出。


内容的提问来源于stack exchange,提问作者Pritish Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 15:25:32