Azure Cache for Redis随机抛出WRONGPASS错误的排查与解决咨询
大家好,我最近碰到了一个棘手的问题,想请社区的大佬们帮忙分析下。我们的Django站点部署在Azure Web App上,搭配Azure Cache for Redis做缓存,全程用托管身份(Managed Identity)来认证Redis连接,没有使用固定的访问密钥。应用是基于Gunicorn多线程模式运行的,为了保证服务可用性,我们自己实现了一个健康检查中间件——不仅会检测Redis和数据库的连通性,还会在检测失败时自动重启Gunicorn做自修复。
最近我们发现这个自修复机制会随机触发,查看日志后定位到是Redis连接时不时抛出WRONGPASS错误。我们的中间件里已经做了3次重试更新Redis凭据的逻辑,但还是会偶尔失败导致服务重启,这给我们的服务稳定性带来了不小的影响。
先贴出我们的健康检查中间件代码:
import logging import os import signal from django.conf import settings from django.core.cache import cache from django.db import connection from django.http import HttpResponse from django_redis import get_redis_connection from tenacity import retry, stop_after_attempt, wait_random_exponential from .azure_helper import RedisCredentials, get_db_password, get_redis_credentials logger = logging.getLogger("pulumi_django_azure.health_check") class HealthCheckMiddleware: def __init__(self, get_response): self.get_response = get_response def _self_heal(self): logger.warning("Self-healing by gracefully restarting Gunicorn.") master_pid = os.getppid() logger.debug("Master PID: %d", master_pid) # Reload a new master with new workers, # since the application is preloaded this is the only safe way for now. os.kill(master_pid, signal.SIGUSR2) # Gracefully shutdown the current workers os.kill(master_pid, signal.SIGTERM) def _test_redis_connection(self) -> bool: try: cache.set("health_check", "test") return True except Exception: return False @retry(stop=stop_after_attempt(3), wait=wait_random_exponential(multiplier=0.5, max=5)) def _update_redis_credentials(self, redis_credentials: RedisCredentials): logger.debug("Updating Redis credentials with password: %s", redis_credentials.password) # Re-authenticate the Redis connection redis_connection = get_redis_connection("default") redis_connection.execute_command("AUTH", redis_credentials.username, redis_credentials.password) settings.CACHES["default"]["OPTIONS"]["PASSWORD"] = redis_credentials.password def __call__(self, request): if request.path == settings.HEALTH_CHECK_PATH: # Update the database credentials if needed if settings.AZURE_DB_PASSWORD: try: current_db_password = settings.DATABASES["default"]["PASSWORD"] new_db_password = get_db_password() if new_db_password != current_db_password: logger.debug("Database password has changed, updating credentials") settings.DATABASES["default"]["PASSWORD"] = new_db_password # Close existing connections to force reconnect with new password connection.close() else: logger.debug("Database password unchanged, keeping existing credentials") except Exception as e: logger.error("Failed to update database credentials: %s", str(e)) self._self_heal() return HttpResponse(status=503) # Update the Redis credentials if needed if settings.AZURE_REDIS_CREDENTIALS: try: current_redis_password = settings.CACHES["default"]["OPTIONS"]["PASSWORD"] redis_credentials = get_redis_credentials() if redis_credentials.password != current_redis_password: logger.debug("Redis password has changed, updating credentials") self._update_redis_credentials(redis_credentials) elif not self._test_redis_connection(): logger.debug("Redis connection check failed, updating credentials") self._update_redis_credentials(redis_credentials) else: logger.debug("Redis password unchanged and connection check passed, keeping existing credentials") except Exception as e: logger.error("Failed to update Redis credentials: %s", str(e)) self._self_heal() return HttpResponse(status=503) try: # Test the database connection connection.ensure_connection() logger.debug("Database connection check passed") # Test the Redis connection cache.set("health_check", "test") logger.debug("Redis connection check passed") return HttpResponse("OK") except Exception as e: logger.error("Health check failed with unexpected error: %s", str(e)) self._self_heal() return HttpResponse(status=503) return self.get_response(request)
从日志里能看到我们确实在更新Redis凭据(其实是托管身份生成的token),但之后还是会出现认证失败:
2025-08-12 13:56:54,647 DEBUG [p:1293] [t:123417039960960] Updating Redis credentials with password: eyJ0eXAiOiJKV1QiLCJhbGciOiJSUzI1NiIsIng1dCI6IkpZaEFjVFBNWl9MWDZEQmxPV1E3SG4wTmVYRSIsImtpZCI6IkpZaEFjVFBNWl9MWDZEQmxPV1E3SG4wTmVYRSJ9.eyJhdWQiOiJodHRwczovL3JlZGlzLmF6dXJlLmNvbSIsImlzcyI6Imh0dHBzOi8vc3RzLndpbmRvd3MubmV0L2UyOThmYWJhLTdjNGEtNGQzZS1iZGQ2LTg3Z...
我现在有几个疑问点想请教大家:
- Gunicorn是多进程/多线程模式,我们在健康检查线程里更新了Redis连接的凭据,但会不会其他线程/进程还在复用旧的连接池连接?
- 我们更新了
settings.CACHES["default"]["OPTIONS"]["PASSWORD"],但Django Redis的缓存后端会不会没有全局更新所有连接池里的连接?有没有办法强制刷新整个连接池? - 目前用
SIGUSR2 + SIGTERM重启Gunicorn的自修复方式是否合理?有没有更优雅的方式更新Redis凭据而不需要重启整个服务? - 托管身份生成的Redis token过期时间是多久?我们的更新逻辑有没有可能在token未过期时重复更新,或者过期后没及时触发更新?
希望有过类似经验的朋友能给点思路,谢谢大家!
内容来源于stack exchange

