Ubuntu下etcd.service启动失败报status=1/FAILURE故障求助
etcd服务启动失败排查方案
故障背景
完成etcd.service配置文件编写后,执行systemctl start etcd.service启动服务时异常退出,服务反复自动重启,启动失败。
现场信息
服务状态查询输出
etcd.service - ve489 etcd service Loaded: loaded (/lib/systemd/system/etcd.service; enabled; vendor preset: enabled) Active: activating (auto-restart) (Result: exit-code) since Sat 2022-06-18 20:59:50 PDT; 3s ago Docs: https://github.com/etcd-io/etcd Process: 13415 ExecStart=/usr/local/bin/etcd --name ubuntu20 --data-dir /var/lib/etcd --initial-advertise-peer-urls http://192.168.159.128:2380 --listen-peer-urls http://192.168.159.128:2380 --listen-client-urls http://192.168.159.128:2379,http://127.0.0.1:2379 --advertise-client-urls http://192.168.159.128:2379 --initial-cluster-token etcd-cluster-1 --initial-cluster etcd-1=http://192.168.159.128:2380,etcd-2=http://192.168.159.129:2380 --initial-cluster-state new --heartbeat-interval 1000 --election-timeout 5000 (code=exited, status=1/FAILURE) Main PID: 13415 (code=exited, status=1/FAILURE)
现有etcd.service配置内容
[Unit] Description=ve489 etcd service Documentation=https://github.com/etcd-io/etcd [Service] User=root Type=notify ExecStart=/usr/local/bin/etcd \ --name ubuntu20 \ --data-dir /var/lib/etcd \ --initial-advertise-peer-urls http://192.168.159.128:2380 \ --listen-peer-urls http://192.168.159.128:2380 \ --listen-client-urls http://192.168.159.128:2379,http://127.0.0.1:2379 \ --advertise-client-urls http://192.168.159.128:2379 \ --initial-cluster-token etcd-cluster-1 \ --initial-cluster etcd-1=http://192.168.159.128:2380,etcd-2=http://192.168.159.129:2380 \ --initial-cluster-state new \ --heartbeat-interval 1000 \ --election-timeout 5000 Restart=on-failure RestartSec=5 [Install] WantedBy=multi-user.target
系统日志查询输出
Jun 18 21:03:57 ubuntu systemd[1]: etcd.service: Main process exited, code=exited, status=1/FAILURE -- Subject: Unit process exited -- Defined-By: systemd -- Support: http://www.ubuntu.com/support -- -- An ExecStart= process belonging to unit etcd.service has exited. -- -- The process' exit code is 'exited' and its exit status is 1. Jun 18 21:03:57 ubuntu systemd[1]: etcd.service: Failed with result 'exit-code'. -- Subject: Unit failed -- Defined-By: systemd -- Support: http://www.ubuntu.com/support -- -- The unit etcd.service has entered the 'failed' state with result 'exit-code'. Jun 18 21:03:57 ubuntu systemd[1]: Failed to start ve489 etcd service. -- Subject: A start job for unit etcd.service has failed -- Defined-By: systemd -- Support: http://www.ubuntu.com/support -- -- A start job for unit etcd.service has finished with a failure. -- -- The job identifier is 38547 and the job result is failed.
根因定位
- 核心配置错误:配置中
--name参数指定当前节点名称为ubuntu20,但--initial-cluster定义的集群成员列表里仅存在etcd-1、etcd-2两个节点名,无匹配当前节点的条目,etcd启动时校验集群成员映射关系失败,直接退出返回状态码1。 - 潜在触发条件:全新部署双节点etcd集群时,若仅启动单节点,另一节点未同步启动,当前节点无法完成集群leader选举,达到选举超时时间后也会触发退出;若
/var/lib/etcd数据目录存在历史残留数据、目录权限与配置中启动用户不匹配,同样会导致启动失败。
修复步骤
- 修正节点名称配置:当前节点IP为192.168.159.128,对应
--initial-cluster中定义的etcd-1节点,将配置文件中--name ubuntu20修改为--name etcd-1,保证节点名和集群列表中的映射键完全一致。 - 清理异常数据(仅全新部署场景操作,存量数据场景跳过):执行
rm -rf /var/lib/etcd/*清除目录下残留的旧集群数据,再执行chown -R root:root /var/lib/etcd确认目录权限和配置中启动用户(root)匹配。 - 重载systemd配置:执行
systemctl daemon-reload加载修改后的服务文件。 - 同步配置另一节点:登录192.168.159.129节点,确认其etcd配置中
--name设置为etcd-2,其余集群token、端口、集群列表参数和当前节点完全一致,完成数据目录权限配置后重载systemd配置。 - 启动集群:两个节点同时执行
systemctl start etcd.service启动服务。 - 验证状态:执行
systemctl status etcd.service确认服务状态为active (running),可通过etcdctl member list命令查询集群成员列表,确认两个节点均正常加入集群。
内容的提问来源于stack exchange,提问作者Naughtybeds
相关产品推荐
相关产品推荐

