DSVM Ubuntu环境下Apache占用GPU内存致预测显存耗尽求助
Hey there, let's dig into this GPU memory issue you're facing—those Apache processes hogging 10GB of your K80's 12GB显存 are definitely the culprit. Here's a step-by-step breakdown of how to fix this:
1. Diagnose the Root Cause
From your nvidia-smi output, /usr/sbin/apache2 is eating up 10860MiB of GPU memory even when there are no active tasks. This happens because Apache's default process model (usually prefork MPM) spins up multiple child processes, and if your Flask app loads the Keras model on process startup, every Apache child process will load its own copy of the model—duplicating GPU memory usage.
2. Tweak Apache's MPM Configuration
First, let's control how many Apache processes are running to reduce redundant model loads:
- Check your current MPM mode:
apache2ctl -M | grep mpm - If you're using
mpm_prefork(the default for many setups), edit its config file (typically/etc/apache2/mods-available/mpm_prefork.conf) to lower process limits:
Restart Apache after changes:<IfModule mpm_prefork_module> StartServers 2 MinSpareServers 2 MaxSpareServers 5 MaxRequestWorkers 10 MaxConnectionsPerChild 1000 </IfModule>systemctl restart apache2 - Alternatively, switch to
mpm_event(a more efficient, thread-based MPM) to reduce overall process count:a2dismod mpm_prefork && a2enmod mpm_event && systemctl restart apache2
3. Lazy-Load Your Keras Model (or Separate Prediction Logic)
Loading the model once per Apache process is still wasteful—here's how to optimize:
- Lazy-load the model: Only load it when the first prediction request comes in, instead of on process startup. Example in your Flask app:
import keras from flask import Flask app = Flask(__name__) model = None def load_model(): global model if model is None: # Load model only once per process model = keras.models.load_model("your_model.h5") return model @app.route("/predict", methods=["POST"]) def predict(): model = load_model() # Run prediction logic here return {"prediction": "result"} - Better yet: Separate prediction from HTTP handling: Use a message queue (like Redis + Celery) to offload prediction tasks to a dedicated worker process. This way, only the Celery worker loads the model once, and Apache/Flask just handles HTTP requests without touching the GPU.
4. Configure TensorFlow/Keras to Use GPU Memory Efficiently
Make sure TensorFlow isn't grabbing more GPU memory than it needs:
import tensorflow as tf from keras.backend.tensorflow_backend import set_session # Enable dynamic GPU memory allocation config = tf.ConfigProto() config.gpu_options.allow_growth = True set_session(tf.Session(config=config))
This tells TensorFlow to only allocate GPU memory as needed, instead of reserving the entire GPU upfront.
5. Consider Switching to a Lighter WSGI Server
Apache can be overkill for Flask apps. Tools like Gunicorn or uWSGI let you better control process/thread counts, reducing redundant model loads:
- Start Flask with Gunicorn (2 worker processes, 4 threads each):
This way, only 2 copies of the model are loaded into GPU memory instead of dozens of Apache child processes.gunicorn --workers 2 --threads 4 your_app_module:app
Verify the Fix
After making changes, run nvidia-smi again to check if Apache's GPU memory usage drops. You should see a significant reduction in idle GPU memory usage, leaving enough space for your prediction batches.
内容的提问来源于stack exchange,提问作者drhoffma

