Celery+RabbitMQ+Redis架构下外部Worker返回结果的2分钟延迟与Socket超时问题排查求助
Let’s walk through your issues step by step, starting with the initial 2-minute delay and then the subsequent timeout error, plus share architecture tweaks to prevent these problems long-term.
Root Cause Analysis
Initial 120-Second Delay (Idle Periods)
The delay happens because:
- Azure’s network infrastructure (NSGs, load balancers, or SNAT) drops idle TCP connections after ~30 minutes by default.
- Your external Workers are running in Azure Container Instances (ACI), which are on a separate network from your Redis VM. When the connection between ACI Worker and Redis goes idle for 30 mins, it gets terminated.
- Redis’s default
redis_socket_timeoutis 120 seconds—so when the Worker tries to send a result back, it waits the full 120s before detecting the dead connection and retrying (sinceredis_retry_on_timeoutwas enabled initially). - The
local_workerdoesn’t hit this issue because it’s on the same VM network as Redis, so idle connections aren’t dropped by external network devices.
Post-Configuration Timeout Error
When you removed redis_retry_on_timeout = True and added redis_socket_keepalive = True, you lost the retry mechanism for dropped connections. The socket_keepalive helps maintain connections, but if it wasn’t configured with aggressive enough parameters to beat Azure’s idle timeout, the connection still gets dropped—and now there’s no retry to recover, leading to redis.exceptions.TimeoutError.
Actionable Fixes
1. Tune Redis & Celery Connection Settings
Update your Celery config to balance keepalive, timeouts, and retries. Here’s the adjusted config:
import os import warnings from pathlib import Path # Result backend use Redis result_backend_host = os.getenv('REDIS_HOST', 'localhost') result_backend_pass = os.getenv('REDIS_PASS', 'password') result_backend = 'redis://:{password}@{host}:6379/0'.format(password=result_backend_pass, host=result_backend_host) # Critical fixes for connection stability redis_retry_on_timeout = True # Re-enable retries for dropped connections redis_socket_timeout = 10 # Shorten timeout to fail fast and retry quickly redis_socket_keepalive = True # Configure aggressive keepalive to prevent Azure from dropping idle connections redis_socket_keepalive_options = { 'TCP_KEEPIDLE': 60, # Start sending keepalive probes after 60s of idle 'TCP_KEEPINTVL': 10, # Send probes every 10s 'TCP_KEEPCNT': 5 # Drop connection after 5 failed probes } # Broker use RabbitMQ rabbitmq_user = os.getenv('RABBITMQ_DEFAULT_USER', 'guest') rabbitmq_pass = os.getenv('RABBITMQ_DEFAULT_PASS', 'guest') rabbitmq_host = os.getenv('RABBITMQ_HOST', 'localhost') broker_url = 'amqp://{user}:{password}@{host}:5672//'.format(user=rabbitmq_user, password=rabbitmq_pass, host=rabbitmq_host) include = ['app.worker.tasks', 'app.dashboard.example1', 'app.dashboard.example2'] # Task events worker_send_task_events = True task_send_sent_event = True
Also, update your Redis container in docker-compose.yml to enable TCP keepalive from the server side:
redis: image: redis:6.2.5 restart: always command: ["redis-server", "--requirepass", "${RABBITMQ_DEFAULT_PASS:-password}", "--tcp-keepalive", "300"] # Send keepalive every 5 mins ports: - 6379:6379 networks: - traefik-public
2. Azure Network Adjustments
- Check VM NSG Rules: Ensure your NSG doesn’t have a custom TCP idle timeout shorter than your keepalive settings. Azure’s default NSG idle timeout is 4 minutes, which our
TCP_KEEPIDLEof 60s will beat. - ACI Network Configuration: If using a virtual network for ACI, confirm that the network peer between your VM’s VNet and ACI’s VNet doesn’t have connection timeout restrictions.
- Avoid Public Port Exposure: Consider moving Redis and RabbitMQ off public ports. Instead, use Azure VNet peering between your VM’s VNet and ACI’s VNet to keep traffic private—this reduces attack surface and avoids issues with public network idle timeouts.
3. Worker Optimization
- Update Redis Client: Ensure your Worker’s
redis-pypackage is up to date (v4.x+ has better keepalive support). - Task Result Handling: For long-running tasks, consider using
ignore_result=Trueif you don’t need to track results, but since you’re using Flower, this might not be feasible. Alternatively, use Celery’sresult_expiresto clean up old results automatically.
Architecture Optimization Suggestions
- Use Azure Redis Cache Instead of Self-Hosted: Managed Azure Redis Cache handles network stability, keepalive, and scaling out of the box. It also has built-in monitoring for connection issues.
- Worker Auto-Scaling: Use Azure Kubernetes Service (AKS) instead of ACI for Workers—AKS lets you auto-scale based on queue length, and you can configure pod lifecycle hooks to clean up connections properly.
- Isolate Services: Move RabbitMQ and Redis to separate VMs or managed services to reduce resource contention on your app VM.
- Monitoring: Add Azure Monitor or Prometheus/Grafana to track Redis connection metrics (like
connected_clients,rejected_connections) and Celery task latency—this helps catch issues before they impact users.
内容的提问来源于stack exchange,提问作者Charles

