Python应用Datadog中冗余Gunicorn日志无法过滤的解决问询
问题
我的Python服务每次运行时都会生成以下日志:
"[2024-05-13 05:58:21 +0000] [67] [INFO] Booting worker with pid: 67" "[2024-05-13 05:58:21 +0000] [1] [INFO] Using worker: sync" "[2024-05-13 05:58:21 +0000] [1] [INFO] Listening at: http://0.0.0.0:5000,http://0.0.0.0:8080 (1)" "[2024-05-13 05:58:21 +0000] [1] [INFO] Starting gunicorn 21.2.0" "[2024-05-13 05:58:21 +0000] [67] [INFO] Worker exiting (pid: 67)" "[2024-05-13 05:58:21 +0000] [1] [INFO] Shutting down: Master" "[2024-05-13 05:58:21 +0000] [1] [INFO] Handling signal: term"
我尝试用以下自定义logging过滤器代码禁用这些日志,但并未生效:
import logging import re class CustomFilter(logging.Filter): def __init__(self, blocked_patterns): super().__init__() self.blocked_patterns = blocked_patterns def filter(self, record): # Check if the log message matches any of the blocked patterns for pattern in self.blocked_patterns: if re.match(pattern, record.getMessage()): # If a blocked pattern is found, return False to filter out the log record return False # If no blocked pattern matches, allow the log record to pass through return True # Configure logging logging.basicConfig(level=logging.INFO) # Define the list of patterns to be blocked blocked_patterns = [ r".*\[INFO\] Booting worker with pid: \d+", r".*\[INFO\] Using worker: .*", r".*\[INFO\] Listening at: http://0\.0\.0\.0:5000,http://0\.0\.0\.0:8080 \(1\)", r".*\[INFO\] Starting gunicorn \d+\.\d+\.\d+", r".*\[INFO\] Worker exiting \(pid: \d+\)", r".*\[INFO\] Shutting down: Master", r".*\[INFO\] Handling signal: term", r".*http://metadata\.google\.internal/computeMetadata/.*" ] # Create a logger logger = logging.getLogger() # Add the custom filter to the logger logger.addFilter(CustomFilter(blocked_patterns))
解决方案
核心原因
Gunicorn拥有独立的日志系统,它的日志输出不会经过你配置的根logger过滤器。Gunicorn会创建专属的logger实例(如gunicorn.error、gunicorn.info等),你的过滤器仅作用于根logger,对这些专用logger无效。
解决步骤
给Gunicorn的logger添加过滤器
直接获取Gunicorn相关的logger实例,为它们添加自定义过滤器:filter_instance = CustomFilter(blocked_patterns) # 遍历Gunicorn的所有相关logger并添加过滤器 for logger_name in ['gunicorn', 'gunicorn.access', 'gunicorn.error', 'gunicorn.info']: gunicorn_logger = logging.getLogger(logger_name) gunicorn_logger.addFilter(filter_instance)简化正则匹配逻辑(可选)
record.getMessage()返回的是原始日志消息内容,不包含时间戳、进程ID等前缀。比如实际日志消息是"Booting worker with pid: 67",而非带时间戳的完整字符串。可以简化正则表达式,仅匹配消息部分:blocked_patterns = [ r"Booting worker with pid: \d+", r"Using worker: .*", r"Listening at: http://0\.0\.0\.0:5000,http://0\.0\.0\.0:8080 \(1\)", r"Starting gunicorn \d+\.\d+\.\d+", r"Worker exiting \(pid: \d+\)", r"Shutting down: Master", r"Handling signal: term", r".*http://metadata\.google\.internal/computeMetadata/.*" ]这样匹配更精准,也能避免前缀格式变化导致的匹配失败。
通过Gunicorn启动参数直接禁用(另一种思路)
如果不需要Gunicorn的所有启动/关闭日志,可以在启动时通过命令行参数调整日志级别:gunicorn --log-level warning your_app:app将日志级别设为
warning后,所有INFO级别的启动日志会被自动过滤。
内容的提问来源于stack exchange,提问作者DeveloperKeycloak
相关产品推荐
相关产品推荐

