Socket.io+NGINX负载均衡配置验证及消息推送丢包问题排查求助
Hey there, let's tackle your problem head-on—first verifying if your NGINX load balancing is actually doing its job, then digging into why only 40% of your users are receiving Socket.io push events. Here's a step-by-step breakdown:
First, we need to confirm traffic is being distributed across your upstream ports as expected:
- Check NGINX Access Logs
Run this command to watch real-time traffic:
Look at the upstream server addresses in the logs—you should see requests hitting all the ports defined in yourtail -f /var/log/nginx/access.logupstreamblock, not just one. If all traffic is going to a single port, your load balancing isn't working. - Add Custom Server Identifiers
Modify each of your Node.js instances to return a custom header with their listening port, like:
Then run repeated curl requests to your domain:app.get('/', (req, res) => { res.setHeader('X-Server-Port', process.env.PORT); res.send('Running'); });
Thecurl -I your-game-domain.comX-Server-Portvalue should rotate across your upstream ports if load balancing is active. - Validate NGINX Config
Make sure your config has no syntax errors and is loaded:
Ifnginx -t systemctl restart nginxnginx -tthrows errors, fix those first—invalid configs won't apply your load balancing rules.
This is almost certainly related to how Socket.io handles state and cross-instance communication. Let's go through the most common issues:
Missing Sticky Sessions (IP Hash)
Socket.io connections are stateful—each user's connection is tied to a specific Node.js instance. If NGINX doesn't use sticky sessions, a user might connect to instance A, but your push event is triggered on instance B. Instance B has no knowledge of that user's connection, so the event never reaches them.Check your
upstreamblock for theip_hashdirective—it should look like this:upstream node_game_servers { ip_hash; # This ensures a user stays connected to the same instance server 127.0.0.1:3000; server 127.0.0.1:3001; server 127.0.0.1:3002; }Without
ip_hash, traffic is distributed randomly, and only users connected to the instance where the push is triggered will get the event (which aligns with your 40% success rate—likely the size of one instance's user pool).Incorrect WebSocket Configuration in NGINX
Socket.io falls back to long polling if WebSocket isn't properly configured in NGINX. Even worse, misconfigured headers can break connections entirely. Ensure yourlocationblock has these settings:location / { proxy_pass http://node_game_servers; proxy_http_version 1.1; proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; proxy_set_header Host $host; proxy_set_header X-Real-IP $remote_addr; # Keep long-lived connections alive proxy_connect_timeout 7d; proxy_send_timeout 7d; proxy_read_timeout 7d; }The
UpgradeandConnectionheaders are critical for maintaining WebSocket connections.No Cross-Instance Event Synchronization
Even with sticky sessions, if you trigger a push event (likeio.emit()) on one Node.js instance, it only broadcasts to users connected to that instance. To send events to all users across all instances, you need a Socket.io adapter that syncs events between servers.The most common solution is the Redis adapter. Install and configure it in your Node.js code:
const { createAdapter } = require('@socket.io/redis-adapter'); const { createClient } = require('redis'); const pubClient = createClient({ url: 'redis://localhost:6379' }); const subClient = pubClient.duplicate(); Promise.all([pubClient.connect(), subClient.connect()]).then(() => { io.adapter(createAdapter(pubClient, subClient)); });This ensures every push event is relayed to all Node.js instances, so all connected users receive it.
Insufficient NGINX Connection Limits
NGINX's defaultworker_connectionsis often too low for 10K concurrent users. Check youreventsblock:events { worker_connections 10000; # Adjust based on your server capacity use epoll; multi_accept on; }If this number is too small, NGINX will reject new connections, leaving users unable to establish a Socket.io session at all.
Check Node.js Instance Health
Use tools likepm2 list(if you're using PM2 to manage instances) to confirm all your upstream Node.js servers are running and not overloaded. If an instance crashes, NGINX will stop sending traffic to it, but any existing connections on that instance will drop, and those users won't receive future pushes.
Even without seeing your config, these are the critical lines to verify:
- Does your
upstreamblock includeip_hash? - Does your
locationblock set theUpgradeandConnectionheaders for WebSocket? - Are all upstream ports in the
upstreamblock mapped to running Node.js instances? - Is
worker_connectionsset high enough to handle 10K concurrent users?
If you share more details from your NGINX config, I can help narrow things down further, but based on your symptoms, sticky sessions and cross-instance event sync are the most likely fixes.
内容的提问来源于stack exchange,提问作者Sorav Garg

