K8s环境下JupyterHub连接Jupyter Enterprise Gateway出现连接超时问题求助
K8s环境下JupyterHub连接Jupyter Enterprise Gateway出现连接超时问题求助
我目前在K8s 1.25集群上部署了JupyterHub 3.2.1和Jupyter Enterprise Gateway(JEG)3.2.1,但是在尝试连接JEG以列出并启动远程内核时,遇到了连接超时的错误,相关日志如下:
Z [I 2024-01-10 13:17:31.189 ServerApp] 200 GET /user/jovyan/api/terminals?1704892657311 (jovyan@127.0.0.1) 0.72ms 2024-01-10T13:18:22.192528667Z [I 2024-01-10 13:18:22.192 ServerApp] 200 GET /user/jovyan/api/me?1704892708321 (jovyan@127.0.0.1) 0.94ms 2024-01-10T13:18:22.764841954Z [I 2024-01-10 13:18:22.764 ServerApp] 101 GET /user/jovyan/api/events/subscribe?token=[secret] (jovyan@127.0.0.1) 0.97ms 2024-01-10T13:19:29.995576942Z [E 2024-01-10 13:19:29.995 ServerApp] Bad message (TypeError('not all arguments converted during string formatting')): {'name': 'ServerApp', 'msg': 'Exception while trying to launch kernel via Gateway URL http://enterprise-gateway.enterprise-gateway.svc.cluster.local:8888 , [Errno 110] Connection timed out', 'args': (TimeoutError(110, 'Connection timed out'),), 'levelname': 'ERROR', 'levelno': 40, 'pathname': '/usr/local/lib/python3.11/site-packages/jupyter_server/gateway/gateway_client.py', 'filename': 'gateway_client.py', 'module': 'gateway_client', 'exc_info': None, 'exc_text': None, 'stack_info': None, 'lineno': 812, 'funcName': 'gateway_request', 'created': 1704892769.9951384, 'msecs': 995.0, 'relativeCreated': 148853.8899421692, 'thread': 140580579428160, 'threadName': 'MainThread', 'processName': 'MainProcess', 'process': 7} 2024-01-10T13:19:29.998097118Z [E 2024-01-10 13:19:29.995 ServerApp] Uncaught exception GET /user/jovyan/api/kernels?1704892645918 (127.0.0.1) 2024-01-10T13:19:29.998121127Z HTTPServerRequest(protocol='https', host='jupyter.xyz.com', method='GET', uri='/user/jovyan/api/kernels?1704892645918', version='HTTP/1.1', remote_ip='127.0.0.1') 2024-01-10T13:19:29.998125801Z Traceback (most recent call last): 2024-01-10T13:19:29.998128775Z File "/usr/local/lib/python3.11/site-packages/tornado/web.py", line 1786, in _execute 2024-01-10T13:19:29.998139181Z result = await result 2024-01-10T13:19:29.998142001Z ^^^^^^^^^^^^
针对这个连接超时问题,我整理了几个实用的排查方向,你可以逐一尝试:
- 验证JEG服务的直接连通性:进入JupyterHub的Pod内部,执行命令
curl http://enterprise-gateway.enterprise-gateway.svc.cluster.local:8888,测试能否正常访问JEG服务。如果连不通,说明网络层面存在阻塞。 - 检查JEG Pod的运行状态:通过
kubectl get pods -n enterprise-gateway查看JEG Pod是否处于Running状态,再用kubectl logs <你的JEG Pod名称> -n enterprise-gateway查看JEG的启动日志,确认服务是否正常监听8888端口。 - 排查NetworkPolicy限制:如果你的K8s集群启用了网络策略,要确认JupyterHub所在Namespace的Pod是否被允许访问JEG所在Namespace的8888端口,有没有对应的允许规则。
- 核对JupyterHub配置:检查JupyterHub的配置中,关于Gateway的URL设置是否正确,比如
c.GatewayClient.url是否准确指向http://enterprise-gateway.enterprise-gateway.svc.cluster.local:8888,有没有拼写错误或者端口写错的情况。 - 测试K8s DNS解析:在JupyterHub Pod里执行
nslookup enterprise-gateway.enterprise-gateway.svc.cluster.local,看看能不能正常解析到JEG服务的ClusterIP。DNS解析失败也会导致连接超时。
备注:内容来源于stack exchange,提问作者AniketGole
相关产品推荐
相关产品推荐

