You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spire Server无法连接PostgreSQL数据库问题求助

问题:Spire Server无法连接同命名空间的PostgreSQL服务

我基于K8s部署Spire Server,初始使用sqlite作为后端时,Spire Server和Agent Pod均正常运行。随后修改server-configmap.yaml切换为PostgreSQL,配置如下:

plugin_data {
    database_type = "postgres"
    connection_string = "dbname=postgres user=postgres password=Jj8Rt9tbyc host=my-postgresql port=5432"
}

PostgreSQL服务部署在同一spire命名空间,状态正常:

$ kubectl get svc my-postgresql -n spire
NAME            TYPE        CLUSTER-IP      EXTERNAL-IP   PORT(S)    AGE
my-postgresql   ClusterIP   10.96.127.179   <none>        5432/TCP   4d2h

$ kubectl get endpoints my-postgresql -n spire
NAME            ENDPOINTS           AGE
my-postgresql   10.244.0.201:5432   4d2h

但Spire Server启动失败,日志报错DNS解析失败:

$ k logs spire-server-0 
time="2024-09-13T10:00:42Z" level=warning msg="Current umask 0022 is too permissive; setting umask 0027"
time="2024-09-13T10:00:42Z" level=info msg=Configured admin_ids="[]" data_dir=/run/spire/data
time="2024-09-13T10:00:42Z" level=info msg="Opening SQL database" db_type=postgres subsystem_name=sql
time="2024-09-13T10:00:47Z" level=error msg="Fatal run error" error="datastore-sql: dial tcp: lookup my-postgresql: Try again"
time="2024-09-13T10:00:47Z" level=error msg="Server crashed" error="datastore-sql: dial tcp: lookup my-postgresql: Try again"

尝试在Spire Server容器内执行nslookup失败(容器无该工具):

kubectl exec -it spire-server-0 -- nslookup my-postgresql
error: Internal error occurred: error executing command in container: failed to exec in container: failed to start exec "0089c4fa4e670cc624f5e24d00b27b06356f936345f19c70cf2545e1cf18a78d": OCI runtime exec failed: exec failed: unable to start container process: exec: "nslookup": executable file not found in $PATH: unknown

但用busybox调试容器可以正常解析PostgreSQL服务名:

$ kubectl run -it --rm debug2 --image=busybox --restart=Never -- sh
If you don't see a command prompt, try pressing enter.
/ # 
/ # nslookup my-postgresql
Server:         10.96.0.10
Address:        10.96.0.10:53

排查解决方案

1. 绕过DNS,直接使用Cluster IP测试

修改server-configmap.yaml中的connection_string,将host=my-postgresql替换为PostgreSQL的Cluster IP(10.96.127.179):

connection_string = "dbname=postgres user=postgres password=Jj8Rt9tbyc host=10.96.127.179 port=5432"

更新ConfigMap后重启Spire Server Pod:

kubectl apply -f server-configmap.yaml -n spire
kubectl delete pod spire-server-0 -n spire

若能正常启动,说明问题出在DNS解析环节。

2. 检查Spire Server Pod的DNS配置

查看Pod的DNS相关配置,确认nameserver和search域是否正常:

kubectl describe pod spire-server-0 -n spire | grep -A10 DNS

正常情况下,nameserver应为K8s集群DNS的Cluster IP(通常是10.96.0.10),search域应包含spire.svc.cluster.local、svc.cluster.local等。

3. 验证容器内DNS配置与连通性

即使没有nslookup,可查看容器内的resolv.conf文件:

kubectl exec -it spire-server-0 -- cat /etc/resolv.conf

若容器有ping命令,直接测试Cluster IP的连通性:

kubectl exec -it spire-server-0 -- ping 10.96.127.179

也可尝试用telnet(若存在)测试端口:

kubectl exec -it spire-server-0 -- telnet 10.96.127.179 5432

4. 检查PostgreSQL的访问控制策略

进入PostgreSQL容器,查看pg_hba.conf是否允许Spire Server所在网段的连接:

kubectl exec -it <postgres-pod-name> -n spire -- cat /var/lib/postgresql/data/pg_hba.conf

确保存在类似规则(根据Spire Pod的网段调整,比如10.244.0.0/16):

host    all             all             10.244.0.0/16          scram-sha-256

修改后需重启PostgreSQL Pod生效。

5. 验证PostgreSQL数据库与用户权限

进入PostgreSQL容器,确认目标数据库和用户存在且有权限:

kubectl exec -it <postgres-pod-name> -n spire -- psql -U postgres

执行以下命令验证:

-- 查看数据库列表
\l
-- 查看用户列表
\du
-- 测试用户登录
SELECT current_user;

内容的提问来源于stack exchange,提问作者Rahul Satal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 07:59:51