Docker Compose部署Swarm时getsockopt权限异常问题排查
Let's break down why your proxy service is failing and fix it step by step:
Root Cause Analysis
That permission error typically pops up when HAProxy tries to run network-related operations (like your option tcp-check config) without sufficient privileges. While you added cap_add in your Compose file, there are a few critical issues preventing those permissions from working as expected, plus some misconfigurations in your HAProxy backend setup.
1. Duplicate volumes Entry in Your Compose File
Look at your proxy service definition: you've got two separate volumes blocks. YAML overwrites duplicate keys, so your first volume mount (/var/run/docker.sock:/var/run/docker.sock) is being ignored. This doesn't just break haproxy-certbot's ability to auto-discover containers—it also indirectly causes permission issues by missing critical socket access.
2. Incorrect Backend Port in HAProxy Config
Your front container exposes port 3000 internally (you mapped it to host port 80 with ports: - '80:3000'), but your HAProxy config tries to connect to front:80. This invalid port check will fail, and combined with the volume issue, amplifies the permission error you're seeing.
3. Redundant Capability Declaration
You added both ALL and NET_ADMIN to cap_add—ALL already includes NET_ADMIN, so this is redundant and can cause unexpected behavior. It's better to only grant the specific capabilities you need.
Fixed Solutions
Step 1: Clean Up Your Docker Compose File
Merge the duplicate volumes blocks, fix capability declarations, and add necessary ports for Let's Encrypt:
version: '3.3' services: back: image: patientplatypus/lowtechback:latest ports: - '5000:5000' deploy: replicas: 3 restart_policy: condition: on-failure max_attempts: 5 window: 120s networks: - web front: image: patientplatypus/lowtechfront:latest ports: - '80:3000' depends_on: - back deploy: replicas: 3 restart_policy: condition: on-failure max_attempts: 5 window: 120s networks: - web proxy: image: nmarus/haproxy-certbot depends_on: - back - front environment: - BALANCE=leastconn ports: - 8080:8080 - 80:80 # Required for Let's Encrypt ACME challenges - 443:443 # Required for HTTPS traffic expose: - "8080" - "3000" - "5000" - "80" - "443" cap_add: - NET_ADMIN # Only grant the necessary network admin capability volumes: - /var/run/docker.sock:/var/run/docker.sock - ./data/config:/config - ./data/letsencrypt:/etc/letsencrypt - ./data/certs:/usr/local/etc/haproxy/certs.d networks: - web deploy: placement: constraints: [node.role == manager] networks: web: driver: overlay
Step 2: Fix HAProxy Backend Configuration
Update ./data/config/haproxy.cfg to use the correct internal port for your front service:
backend my_http_backend mode http balance leastconn option tcp-check option log-health-checks server back back:5000 check port 5000 server front front:3000 check port 3000 # Fixed to use container's internal port
Step 3: Redeploy the Stack
Clean up the old deployment and restart fresh:
# Remove the existing stack docker stack rm prod2 # Clean up unused resources to avoid conflicts docker system prune -f # Redeploy with the fixed config docker stack deploy --compose-file=docker-compose.yaml prod2
Extra Validation Checks
- Ensure the
./datadirectory has proper permissions: runchmod -R 755 ./datato let the proxy container read/write to these volumes. - Verify your node's firewall allows traffic on ports 80, 443, and 8080 (critical for Let's Encrypt and your service access).
- Confirm no network policies in your Swarm cluster are blocking communication between the proxy, front, and back services.
内容的提问来源于stack exchange,提问作者Peter Weyand

