TensorFlow 2.6-2.9大数据集训练时GPU内存不足问题求助
问题:TensorFlow 2.6-2.9大数据集训练GPU内存耗尽
现象
- 线性自编码器在1000条记录的小数据集上训练正常,100万条记录的大数据集上出现GPU内存耗尽错误
- TensorFlow 2.5版本运行正常,2.6-2.9版本崩溃,CPU训练始终正常
- 批大小固定(16),主机内存充足,小数据集仅占用不到800MiB GPU显存
环境信息
- Python 3.9 + Fedora 36
- GPU:Nvidia RTX 2070(8GiB显存)
- 驱动版本515.48.07,CUDA版本11.7
模型代码
def get_model(n_inputs: int) -> models.Model: inp = layers.Input(shape=(n_inputs,)) out = layers.Dense(n_inputs, activation='linear')(inp) m = models.Model(inputs=inp, outputs=out) m.compile(loss='mse', optimizer='adam') m.summary() return m
数据加载代码
def wrap_data(data: np.ndarray) -> tf.data.Dataset: dataset = tf.data.Dataset.from_tensor_slices(data) shuffled = dataset.shuffle(buffer_size=len(data), reshuffle_each_iteration=True) batched = shuffled.batch(16, num_parallel_calls=tf.data.AUTOTUNE, deterministic=False) autoencoder = batched.map(lambda x: (x, x)).prefetch(5) return autoencoder
报错信息
初始运行报错
2022-09-05 15:29:37.525261: W tensorflow/core/framework/cpu_allocator_impl.cc:82] Allocation of 16384000000 exceeds 10% of free system memory. 2022-09-05 15:29:54.002629: W tensorflow/core/common_runtime/bfc_allocator.cc:479] Allocator (GPU_0_bfc) ran out of memory trying to allocate 15.26GiB (rounded to 16384000000)requested by op _EagerConst If the cause is memory fragmentation maybe the environment variable 'TF_GPU_ALLOCATOR=cuda_malloc_async' will improve the situation. Current allocation summary follows. Current allocation summary follows. 2022-09-05 15:29:54.002987: W tensorflow/core/common_runtime/bfc_allocator.cc:491] *_******____________________________________________________________________________________________ Traceback (most recent call last): File "/home/david/[path]/benchmark.py", line 49, in <module> main(parser.parse_args().big) File "/home/david/[path]/benchmark.py", line 40, in main train_data_iterator = wrap_data(train_data) File "/home/david/[path]/benchmark.py", line 33, in wrap_data dataset = tf.data.Dataset.from_tensor_slices(data) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/ops/dataset_ops.py", line 809, in from_tensor_slices return TensorSliceDataset(tensors, name=name) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/ops/dataset_ops.py", line 4551, in __init__ element = structure.normalize_element(element) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/util/structure.py", line 125, in normalize_element ops.convert_to_tensor(t, name="component_%d" % i, dtype=dtype)) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/profiler/trace.py", line 183, in wrapped return func(*args, **kwargs) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/ops.py", line 1640, in convert_to_tensor ret = conversion_func(value, dtype=dtype, name=name, as_ref=as_ref) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/tensor_conversion_registry.py", line 48, in _default_conversion_function return constant_op.constant(value, dtype, name=name) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 267, in constant return _constant_impl(value, dtype, shape, name, verify_shape=False, File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 279, in _constant_impl return _constant_eager_impl(ctx, value, dtype, shape, verify_shape) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 304, in _constant_eager_impl t = convert_to_eager_tensor(value, ctx, dtype) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 102, in convert_to_eager_tensor return ops.EagerTensor(value, ctx.device_name, dtype) tensorflow.python.framework.errors_impl.InternalError: Failed copying input tensor from /job:localhost/replica:0/task:0/device:CPU:0 to /job:localhost/replica:0/task:0/device:GPU:0 in order to run _EagerConst: Dst tensor is not initialized.
指定GPU内存分配器后报错
2022-09-05 15:33:19.542433: W tensorflow/core/framework/cpu_allocator_impl.cc:82] Allocation of 16384000000 exceeds 10% of free system memory. 2022-09-05 15:33:25.973935: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:288] gpu_async_0 cuMemAllocAsync failed to allocate 16384000000 bytes: CUDA error: out of memory (CUDA_ERROR_OUT_OF_MEMORY) Reported by CUDA: Free memory/Total memory: 1115357184/8369799168 2022-09-05 15:33:25.973961: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:293] Stats: Limit: 6221922304 InUse: 67126312 MaxInUse: 201327628 NumAllocs: 13 MaxAllocSize: 67108864 Reserved: 0 PeakReserved: 0 LargestFreeBlock: 0 2022-09-05 15:33:25.973970: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:56] Histogram of current allocation: (allocation_size_in_bytes, nb_allocation_of_that_sizes), ...; 2022-09-05 15:33:25.973974: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 4, 5 2022-09-05 15:33:25.973976: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 8, 2 2022-09-05 15:33:25.973979: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 1028, 1 2022-09-05 15:33:25.973982: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 16384, 1 2022-09-05 15:33:25.973985: E tensorflow/core/common_runtime/gpu/gpu_cudamallocasync_allocator.cc:59] 67108864, 1 Traceback (most recent call last): File "/home/david/[path]/benchmark.py", line 48, in <module> main(parser.parse_args().big) File "/home/david/[path]/benchmark.py", line 40, in main train_data_iterator = wrap_data(train_data) File "/home/david/[path]/benchmark.py", line 33, in wrap_data dataset = tf.data.Dataset.from_tensor_slices(data) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/ops/dataset_ops.py", line 809, in from_tensor_slices return TensorSliceDataset(tensors, name=name) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/ops/dataset_ops.py", line 4551, in __init__ element = structure.normalize_element(element) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/data/util/structure.py", line 125, in normalize_element ops.convert_to_tensor(t, name="component_%d" % i, dtype=dtype)) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/profiler/trace.py", line 183, in wrapped return func(*args, **kwargs) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/ops.py", line 1640, in convert_to_tensor ret = conversion_func(value, dtype=dtype, name=name, as_ref=as_ref) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/tensor_conversion_registry.py", line 48, in _default_conversion_function return constant_op.constant(value, dtype, name=name) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 267, in constant return _constant_impl(value, dtype, shape, name, verify_shape=False, File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 279, in _constant_impl return _constant_eager_impl(ctx, value, dtype, shape, verify_shape) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 304, in _constant_eager_impl t = convert_to_eager_tensor(value, ctx, dtype) File "/home/david/.virtualenvs/ainet/lib64/python3.9/site-packages/tensorflow/python/framework/constant_op.py", line 102, in convert_to_eager_tensor return ops.EagerTensor(value, ctx.device_name, dtype) tensorflow.python.framework.errors_impl.InternalError: Failed copying input tensor from /job:localhost/replica:0/task:0/device:CPU:0 to /job:localhost/replica:0/task:0/device:GPU:0 in order to run _EagerConst: Dst tensor is not initialized.
后续处理
已提交Bug报告
内容的提问来源于stack exchange,提问作者Davidmh
相关产品推荐
相关产品推荐

