You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DSVM Ubuntu环境下Apache占用GPU内存致预测显存耗尽求助

Fixing Apache-Keras GPU Memory Leak on DSVM Ubuntu

Hey there, let's dig into this GPU memory issue you're facing—those Apache processes hogging 10GB of your K80's 12GB显存 are definitely the culprit. Here's a step-by-step breakdown of how to fix this:

1. Diagnose the Root Cause

From your nvidia-smi output, /usr/sbin/apache2 is eating up 10860MiB of GPU memory even when there are no active tasks. This happens because Apache's default process model (usually prefork MPM) spins up multiple child processes, and if your Flask app loads the Keras model on process startup, every Apache child process will load its own copy of the model—duplicating GPU memory usage.

2. Tweak Apache's MPM Configuration

First, let's control how many Apache processes are running to reduce redundant model loads:

  • Check your current MPM mode:
    apache2ctl -M | grep mpm
    
  • If you're using mpm_prefork (the default for many setups), edit its config file (typically /etc/apache2/mods-available/mpm_prefork.conf) to lower process limits:
    <IfModule mpm_prefork_module>
        StartServers             2
        MinSpareServers          2
        MaxSpareServers          5
        MaxRequestWorkers        10
        MaxConnectionsPerChild   1000
    </IfModule>
    
    Restart Apache after changes:
    systemctl restart apache2
    
  • Alternatively, switch to mpm_event (a more efficient, thread-based MPM) to reduce overall process count:
    a2dismod mpm_prefork && a2enmod mpm_event && systemctl restart apache2
    

3. Lazy-Load Your Keras Model (or Separate Prediction Logic)

Loading the model once per Apache process is still wasteful—here's how to optimize:

  • Lazy-load the model: Only load it when the first prediction request comes in, instead of on process startup. Example in your Flask app:
    import keras
    from flask import Flask
    
    app = Flask(__name__)
    model = None
    
    def load_model():
        global model
        if model is None:
            # Load model only once per process
            model = keras.models.load_model("your_model.h5")
        return model
    
    @app.route("/predict", methods=["POST"])
    def predict():
        model = load_model()
        # Run prediction logic here
        return {"prediction": "result"}
    
  • Better yet: Separate prediction from HTTP handling: Use a message queue (like Redis + Celery) to offload prediction tasks to a dedicated worker process. This way, only the Celery worker loads the model once, and Apache/Flask just handles HTTP requests without touching the GPU.

4. Configure TensorFlow/Keras to Use GPU Memory Efficiently

Make sure TensorFlow isn't grabbing more GPU memory than it needs:

import tensorflow as tf
from keras.backend.tensorflow_backend import set_session

# Enable dynamic GPU memory allocation
config = tf.ConfigProto()
config.gpu_options.allow_growth = True
set_session(tf.Session(config=config))

This tells TensorFlow to only allocate GPU memory as needed, instead of reserving the entire GPU upfront.

5. Consider Switching to a Lighter WSGI Server

Apache can be overkill for Flask apps. Tools like Gunicorn or uWSGI let you better control process/thread counts, reducing redundant model loads:

  • Start Flask with Gunicorn (2 worker processes, 4 threads each):
    gunicorn --workers 2 --threads 4 your_app_module:app
    
    This way, only 2 copies of the model are loaded into GPU memory instead of dozens of Apache child processes.

Verify the Fix

After making changes, run nvidia-smi again to check if Apache's GPU memory usage drops. You should see a significant reduction in idle GPU memory usage, leaving enough space for your prediction batches.

内容的提问来源于stack exchange,提问作者drhoffma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:37:22