You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Gunicorn部署的Flask请求,能否限制其CPU/内存占用?

How to Add Hard Resource Limits for Gunicorn/Flask to Defend Against DoS Attacks

Great question—prioritizing defense in depth against attacks like the Billion Laughs XML bomb is absolutely the right call, especially when you want to stop a single malicious request from taking down your entire service. Let’s break down how to approach resource limits, since there are a few key nuances to Gunicorn’s architecture and OS-level controls you’ll need to know.

First: Clarify Gunicorn’s Request Handling Model

You mentioned "each request corresponds to a thread"—this depends on the worker type you’re using:

  • Sync workers (default): Each worker is a single process that handles one request at a time. No threads involved here.
  • Threaded workers: Each worker process spawns multiple threads, with each thread handling one request. Threads in the same worker share the process’s memory space, which is critical for resource limiting.

The key takeaway: Threads can’t be individually restricted with hard memory limits (since they share the parent process’s address space), but you can restrict the worker processes that host those threads. CPU limits are possible but trickier at the thread level—process-level limits are more practical.

Practical Solutions for Defense in Depth

Let’s layer your defenses from application-level up to OS-level, so you have multiple safeguards:

1. Application-Level Hardening (First Line of Defense)

Even with resource limits, stop malicious payloads before they can consume resources:

  • Limit JSON/XML parsing: For Billion Laughs attacks (which target recursive expansion), use parsers that enforce limits. For example:
    • If using JSON, use simplejson with max_depth and max_size parameters instead of Python’s default json module.
    • If handling XML, use defusedxml instead of Python’s built-in xml modules—it blocks recursive entity expansion by default.
  • Set Flask’s MAX_CONTENT_LENGTH: Add app.config['MAX_CONTENT_LENGTH'] = 16 * 1024 * 1024 to reject requests larger than 16MB.
  • Use Gunicorn’s request limits: Enable --limit-request-line 4096 and --limit-request-field_size 8192 to cap header sizes, which blocks some header-based attacks.

2. Gunicorn Worker Process Limits

Since worker processes (even threaded ones) are isolated, you can use OS tools to restrict their resource usage:

  • Use ulimit at startup: When launching Gunicorn, set per-process memory limits with ulimit. For example, to cap each worker’s virtual memory at 512MB:
    ulimit -v 524288 && gunicorn --workers 4 --threads 2 myapp:app
    
    This ensures any worker that exceeds the memory limit gets terminated by the OS.
  • Restart workers periodically: Use --max-requests 1000 to have workers restart after handling 1000 requests. This prevents memory leaks from accumulating over time.

3. OS-Level Hardening with cgroups or systemd

For more granular control (like CPU quotas or strict memory limits), use OS-level tools:

  • If using systemd to manage Gunicorn: Add resource limits directly to your .service file. For example:
    [Service]
    ExecStart=/usr/bin/gunicorn --workers 4 --threads 2 myapp:app
    MemoryLimit=512M  # Hard limit per service (split across workers)
    CPUQuota=50%      # Limit total CPU usage to 50% of one core
    
    Systemd uses cgroups under the hood to enforce these limits, and it will restart any workers that exceed them.
  • Manual cgroup setup: If you’re not using systemd, you can create a cgroup, set memory/CPU limits, and launch Gunicorn workers into it. This is more complex but gives you per-worker control if needed.

Key Takeaway

You’re right that the OS handles most hard resource limiting—Gunicorn itself doesn’t provide thread-level resource controls (and for good reason, since threads share process resources). Instead, focus on:

  • Isolating requests with worker processes (prefork or threaded model)
  • Applying OS-level limits to those workers
  • Layering in application-level checks to block malicious payloads before they can consume resources

This defense-in-depth approach ensures that even if one layer is bypassed, others will stop the attack from affecting other users.

内容的提问来源于stack exchange,提问作者Martin Thoma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:59:44