ActionCable WebSocket连接每3分钟自动断开问题排查求助
WebSocket连接每3分钟自动断开排查求助
问题描述
在开发与生产环境中,WebSocket连接每3分钟左右(精确在2.95至2.98分钟之间)就会断开。已在客户端实现每10秒发送一次ping消息防止闲置,但无论连接活跃度、消息收发量如何,连接总会准时终止。
相关代码
Rails Channel 代码
class SubscriptionIntentionsSelectionChannel < ApplicationCable::Channel def subscribed Rails.logger.info("subscribed") stream_for Project.find(6) end def receive(_data) self.class.broadcast_to(Project.find(6), status: :success, message: "pong") end end
客户端 JavaScript 代码
$('document').ready(() => { channel = App.cable.subscriptions.create( { channel: 'SubscriptionIntentionsSelectionChannel' }, { received(data) { console.log(data); }, rejected() { }, disconnected() { } }); setInterval(() => { channel.send({ data: 'ping' }); }, 10000); });
日志信息
服务器日志
SubscriptionIntentionsSelectionChannel is transmitting the subscription confirmation SubscriptionIntentionsSelectionChannel is streaming from subscription_intentions_selection:Z2lkOi8vYmxhc3QvUHJvamVjdC82 Finished "/cable/" [WebSocket] for 127.0.0.1 at 2023-05-04 11:22:12 +0200 SubscriptionIntentionsSelectionChannel stopped streaming from subscription_intentions_selection:Z2lkOi8vYmxhc3QvUHJvamVjdC82
浏览器开发者控制台日志

已尝试的排查动作
- 将
cable.yml中的适配器从 redis 切换为 async,问题未解决 - 调整
config/puma.rb中的worker_timeout参数,无效
求助问题
有什么方法可以获取连接断开的具体原因?
排查建议
检查反向代理/负载均衡的超时设置
这种精确3分钟的断开,大概率是反向代理(如Nginx、Apache)或负载均衡器的超时配置导致的。比如Nginx默认proxy_read_timeout为180秒(刚好3分钟),完全匹配你的断开时长。- 找到代理配置文件,检查
proxy_connect_timeout、proxy_send_timeout、proxy_read_timeout这几个参数,同时确认WebSocket升级头是否正确配置:proxy_set_header Upgrade $http_upgrade; proxy_set_header Connection "upgrade"; - 暂时延长超时时间(比如设置为300秒),测试是否还会准时断开。
- 找到代理配置文件,检查
开启Action Cable与Puma的详细日志
- 在Rails环境配置文件(如
config/environments/production.rb)中提升Action Cable的日志级别:
这样能捕获到更多连接建立、心跳、断开阶段的细节日志,可能会明确记录断开的触发源。config.action_cable.log_level = :debug - 开启Puma调试日志,在
config/puma.rb中添加:
查看Puma日志中是否有连接关闭的相关提示。stdout_redirect './log/puma.stdout.log', './log/puma.stderr.log', true debug
- 在Rails环境配置文件(如
检查操作系统TCP Keepalive参数
部分系统的TCP Keepalive默认时长接近180秒,可通过命令查看:- Linux:
sysctl net.ipv4.tcp_keepalive_time - macOS:
sysctl net.inet.tcp.keepidle
如果值为180秒左右,可尝试调整该参数,或者在config/cable.yml中开启Action Cable的TCP Keepalive:production: adapter: redis url: <%= ENV.fetch("REDIS_URL") { "redis://localhost:6379/1" } %> tcp_keepalive: true
- Linux:
客户端捕获断开时的错误信息
修改客户端disconnected回调,增加错误信息捕获:disconnected() { const disconnectError = App.cable.connection.disconnectError; console.error('WebSocket断开原因:', disconnectError); }同时监听连接错误事件:
App.cable.connection.addEventListener('error', (error) => { console.error('WebSocket连接错误:', error); });跳过反向代理直接测试连接
如果是生产环境,尝试直接访问服务器的WebSocket端口(如默认28080),绕过Nginx等代理,观察连接是否还会准时断开。如果不再断开,即可确认是代理的超时配置问题。
内容的提问来源于stack exchange,提问作者Kernael
相关产品推荐
相关产品推荐

