新Windows11设备TensorFlow调用CUDA报错:ptxas编译PTX到SASS失败
CUDA与TensorFlow环境配置训练报错求助
环境版本
- CUDA 11.2
- cuDNN 8.9
- TensorFlow 2.10
- Python 3.10
- 显卡:NVIDIA GeForce RTX 4070 Ti(计算能力8.9)
- 系统:Windows 11
配置过程
最初尝试Miniconda安装失败,卸载后改用Python 3.10重新部署环境,已添加系统环境变量:
C:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.2\binC:\Program Files\NVIDIA GPU Computing Toolkit\CUDA\v11.2\libnvvp
完整报错日志
2023-09-13 23:15:01.935221: I tensorflow/core/platform/cpu_feature_guard.cc:193] This TensorFlow binary is optimized with oneAPI Deep Neural Network Library (oneDNN) to use the following CPU instructions in performance-critical operations: AVX AVX2 To enable them in other operations, rebuild TensorFlow with the appropriate compiler flags. 2023-09-13 23:15:02.279467: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1616] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 9392 MB memory: -> device: 0, name: NVIDIA GeForce RTX 4070 Ti, pci bus id: 0000:01:00.0, compute capability: 8.9 Model: "sequential" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= lstm (LSTM) (24, 64) 20736 batch_normalization (BatchN (24, 64) 256 ormalization) dropout (Dropout) (24, 64) 0 dense (Dense) (24, 8) 520 batch_normalization_1 (Batc (24, 8) 32 hNormalization) dropout_1 (Dropout) (24, 8) 0 dense_1 (Dense) (24, 1) 9 ================================================================= Total params: 21,553 Trainable params: 21,409 Non-trainable params: 144 _________________________________________________________________ 2023-09-13 23:15:03.550772: W tensorflow/core/framework/cpu_allocator_impl.cc:82] Allocation of 4608000000 exceeds 10% of free system memory. Epoch 1/10 2023-09-13 23:15:06.027973: I tensorflow/stream_executor/cuda/cuda_dnn.cc:384] Loaded cuDNN version 8905 Could not load symbol cublasGetSmCountTarget from cublas64_11.dll. Error code 127 2023-09-13 23:15:06.342783: I tensorflow/stream_executor/cuda/cuda_blas.cc:1614] TensorFloat-32 will be used for the matrix multiplication. This will only be logged once. 2023-09-13 23:15:06.385428: I tensorflow/compiler/xla/service/service.cc:173] XLA service 0x22167e1f550 initialized for platform CUDA (this does not guarantee that XLA will be used). Devices: 2023-09-13 23:15:06.385546: I tensorflow/compiler/xla/service/service.cc:181] StreamExecutor device (0): NVIDIA GeForce RTX 4070 Ti, Compute Capability 8.9 2023-09-13 23:15:06.389863: I tensorflow/compiler/mlir/tensorflow/utils/dump_mlir_util.cc:268] disabling MLIR crash reproducer, set env var `MLIR_CRASH_REPRODUCER_DIRECTORY` to enable. 2023-09-13 23:15:06.470150: F tensorflow/compiler/xla/service/gpu/nvptx_compiler.cc:453] ptxas returned an error during compilation of ptx to sass: 'INTERNAL: ptxas exited with non-zero error code -1, output: ' If the error message indicates that a file could not be written, please verify that sufficient filesystem space is provided. Process finished with exit code -1073740791 (0xC0000409)
训练代码
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import * from tensorflow.keras.callbacks import ModelCheckpoint from tensorflow.keras.losses import BinaryCrossentropy from tensorflow.keras.optimizers import Adam from tensorflow.keras.optimizers import AdamW import tensorflow as tf import data_formatting import pandas as pd from sklearn.preprocessing import Normalizer data = 'data/' X_train, y_train, X_cv, y_cv, X_test, y_test, init_bias, class_weight = data_formatting.transform(data=data) output_bias = tf.keras.initializers.Constant(init_bias) model1 = Sequential([ InputLayer(batch_input_shape=(24, 300, 16)), LSTM(units=64, stateful=True), tf.keras.layers.BatchNormalization(), tf.keras.layers.Dropout(0.3), Dense(units=8, activation='tanh', kernel_regularizer=tf.keras.regularizers.L2(0.16)), tf.keras.layers.BatchNormalization(), tf.keras.layers.Dropout(0.3), Dense(units=1, activation='sigmoid', bias_initializer=output_bias) ]) model1.summary() cp = ModelCheckpoint('ModelTest/', save_best_only=True) model1.compile(loss=BinaryCrossentropy(), optimizer=AdamW(learning_rate=0.0001), metrics=['accuracy', tf.keras.metrics.Precision(), tf.keras.metrics.Recall()]) model1.fit(X_train, y_train, validation_data=(X_cv, y_cv), epochs=10, batch_size=24, callbacks=[cp], class_weight=class_weight)
已搜索相关报错但未找到有效解决方案,特此求助。
内容的提问来源于stack exchange,提问作者Didlex
相关产品推荐
相关产品推荐

