使用Keras训练CNN时遭遇GPU资源耗尽(OOM)问题求助
Hey there, let's work through this out-of-memory (OOM) issue you're hitting. The GeForce 940MX is a capable entry-level GPU, but it typically only has 2GB or 4GB of VRAM—so we need to dig beyond just shrinking batch size to free up space. Here are actionable steps tailored to your setup:
1. Trim Your GloVE Embedding Matrix (Biggest Win)
The full GloVE.6B.100d file has a vocabulary of ~400k words, but your dataset almost certainly doesn't use all of them. Loading only the words that actually appear in your data will drastically reduce the size of your embedding tensor (the one throwing the OOM error right now).
Here's how to do it:
from keras.preprocessing.text import Tokenizer import numpy as np # First, get the unique words from your dataset tokenizer = Tokenizer() tokenizer.fit_on_texts(your_text_data) word_index = tokenizer.word_index # Load only GloVE vectors for words in your dataset embeddings_index = {} with open("glove.6B.100d.txt", encoding="utf8") as f: for line in f: word, coefs = line.split(maxsplit=1) coefs = np.fromstring(coefs, "f", sep=" ") if word in word_index: embeddings_index[word] = coefs # Build a smaller embedding matrix EMBEDDING_DIM = 100 embedding_matrix = np.zeros((len(word_index) + 1, EMBEDDING_DIM)) for word, i in word_index.items(): embedding_vector = embeddings_index.get(word) if embedding_vector is not None: embedding_matrix[i] = embedding_vector
Also, set trainable=False on your embedding layer (at least initially)—this skips storing gradient information for the embedding weights, saving extra VRAM. You can enable fine-tuning later if needed.
2. Reduce MAX_SEQUENCE_LENGTH
A sequence length of 500 might be overkill for your data. Check the actual length distribution of your text:
text_lengths = [len(text.split()) for text in your_text_data] print(f"95th percentile text length: {np.percentile(text_lengths, 95)}")
Set MAX_SEQUENCE_LENGTH to the 95th or 90th percentile value (e.g., 200 or 300) instead of 500. This cuts down the size of each input sample, reducing the total memory per batch.
3. Lighten Your CNN Model
- Cut filter counts: If you're using layers like
Conv1D(64, 3), try dropping toConv1D(32, 3)first—fewer filters mean fewer parameters and less memory usage. - Shrink dense layers: Replace large fully connected layers (e.g.,
Dense(512)) with smaller ones likeDense(256)orDense(128). - Add Dropout: Not only does it prevent overfitting, but Dropout layers reduce the number of active neurons during training, lowering memory demands.
4. GPU Memory Tweaks
- Enable mixed precision training: This uses float16 tensors for most computations, cutting VRAM usage in half. The 940MX supports this:
import tensorflow as tf tf.keras.mixed_precision.set_global_policy("mixed_float16") - Enable memory growth: Tell TensorFlow to allocate VRAM only as needed, instead of grabbing all available space upfront:
gpus = tf.config.experimental.list_physical_devices("GPU") if gpus: try: for gpu in gpus: tf.config.experimental.set_memory_growth(gpu, True) except RuntimeError as e: print(e)
5. Avoid Loading All Data at Once
Make sure you're using a data generator (like tf.data.Dataset or Keras' Sequence class) to load batches on-the-fly, instead of loading all 65k files into CPU memory at once. This prevents CPU memory overflow from spilling over into GPU memory.
Start with steps 1 and 2—they'll give you the biggest memory savings quickly. If you still hit OOM, work through the rest of the list.
内容的提问来源于stack exchange,提问作者user9329845

