Google Cloud Debian VM升级后浏览器SSH报错无可用认证方法
问题描述
我有一台运行Debian系统的Google Cloud VM实例,托管了WordPress站点。系统升级后初期运行正常,可通过「在浏览器中打开SSH」选项正常连接SSH。
现在尝试使用该功能连接实例时持续重试,检查串口控制台输出无报错信息。
但使用同一密钥可正常通过FTP连接,仅SSH连接存在问题,已确认实例及项目的22端口处于开放状态。
重启VM后串口控制台日志最后几行如下:
Dec 6 09:01:20 localhost sendmail[383]: Starting Mail Transport Agent (MTA): sendmail. Dec 6 09:01:20 localhost systemd[1]: Started LSB: powerful, efficient, and scalable Mail Transport Agent. Dec 6 09:01:21 localhost systemd[1]: Started MariaDB 10.3.31 database server. Dec 6 09:01:21 localhost systemd[1]: Reached target Multi-User System. Dec 6 09:01:21 localhost systemd[1]: Reached target Graphical InterfacDec 6 09:01:21 localhost systemd[1]: Startup finished in 4.063s (kernel) + 9.852s (userspace) = 13.915s. Dec 6 09:01:21 localhost /etc/mysql/debian-start[567]: Upgrading MySQL tables if necessary. Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: /usr/bin/mysql_upgrade: the '--basedir' option is always ignored Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: Looking for 'mysql' as: /usr/bin/mysql Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: Looking for 'mysqlcheck' as: /usr/bin/mysqlcheck Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: Version check failed. Got the following error when calling the 'mysql' command line client Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: ERROR 1045 (28000): Access denied for user 'root'@'localhost' (using password: NO) Dec 6 09:01:21 localhost /etc/mysql/debian-start[570]: FATAL ERROR: Upgrade failed Dec 6 09:01:21 localhost /etc/mysql/debian-start[580]: Checking for insecure root accounts. Dec 6 09:01:21 localhost debian-start[564]: ERROR 1045 (28000): Access denied for user 'root'@'localhost' (using password: NO) Debian GNU/Linux 10 localhost ttyS0 localhost login: Dec 6 09:01:28 localhost systemd[1]: Stopping User Manager for UID 110... Dec 6 09:01:28 localhost systemd[497]: Stopped target Default. Dec 6 09:01:28 localhost systemd[497]: Stopped target Basic System. Dec 6 09:01:28 localhost systemd[497]: Stopped target Timers. Dec 6 09:01:28 localhost systemd[497]: Stopped target Paths. Dec 6 09:01:28 localhost systemd[497]: Stopped target Sockets. Dec 6 09:01:28 localhost systemd[497]: gpg-agent-browser.socket: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Closed GnuPG cryptographic agent and passphrase cache (access for web browsers). Dec 6 09:01:28 localhost systemd[497]: dirmngr.socket: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Closed GnuPG network certificate management daemon. Dec 6 09:01:28 localhost systemd[497]: gpg-agent-ssh.socket: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Closed GnuPG cryptographic agent (ssh-agent emulation). Dec 6 09:01:28 localhost systemd[497]: gpg-agent.socket: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Closed GnuPG cryptographic agent and passphrase cache. Dec 6 09:01:28 localhost systemd[497]: gpg-agent-extra.socket: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Closed GnuPG cryptographic agent and passphrase cache (restricted). Dec 6 09:01:28 localhost systemd[497]: Reached target Shutdown. Dec 6 09:01:28 localhost systemd[497]: systemd-exit.service: Succeeded. Dec 6 09:01:28 localhost systemd[497]: Started Exit the Session. Dec 6 09:01:28 localhost systemd[497]: Reached target Exit the Session. Dec 6 09:01:28 localhost systemd[1]: user@110.service: Succeeded. Dec 6 09:01:28 localhost systemd[1]: Stopped User Manager for UID 110. Dec 6 09:01:28 localhost systemd[1]: Stopping User Runtime Directory /run/user/110... Dec 6 09:01:28 localhost systemd[1]: run-user-110.mount: Succeeded. Dec 6 09:01:28 localhost systemd[1]: user-runtime-dir@110.service: Succeeded. Dec 6 09:01:28 localhost systemd[1]: Stopped User Runtime Directory /run/user/110. Dec 6 09:01:28 localhost systemd[1]: Removed slice User Slice of UID 110.
已尝试的解决方案
- 解决方案1:使用PuTTYGen & Putty
通过PuTTYGen生成密钥,将公钥添加到元数据及实例配置中,已将enable-oslogin设置为FALSE,连接仍报错 - 解决方案2:使用串口连接
尝试通过不同串口连接时卡在连接界面,对应串口控制台日志为空 - 解决方案3:基于磁盘镜像创建新实例
为当前磁盘创建镜像后使用该镜像新建实例,连接新实例时仍存在相同问题 - 解决方案4:在其他设备配置CLI
在新设备上全新安装Google Cloud CLI后尝试连接,同样失败,报错信息一致 - 解决方案5:扩容磁盘空间
将磁盘空间从20GB扩容至35GB,问题仍未解决;通常磁盘空间不足的报错会出现在串口控制台日志中,但本次日志无相关错误
排查修复建议
优先通过启动脚本修复SSH服务
该问题大概率是系统升级后SSH配置被修改、或者SSH服务未正常启动导致,可通过GCP控制台的自定义元数据启动脚本临时注入命令修复:
- 停止当前VM实例
- 编辑实例配置,在自定义元数据中添加键为
startup-script,值为以下内容:
#!/bin/bash # 备份并重置SSH默认配置 cp /etc/ssh/sshd_config /etc/ssh/sshd_config.bak cp /usr/share/openssh/sshd_config /etc/ssh/sshd_config # 临时开启密码登录,修复后可关闭 echo "PasswordAuthentication yes" >> /etc/ssh/sshd_config echo "PubkeyAuthentication yes" >> /etc/ssh/sshd_config # 重启并设置SSH服务开机自启 systemctl restart sshd systemctl enable sshd # 临时添加测试管理员用户,用户名test,密码test123 useradd -m -s /bin/bash test echo "test:test123" | chpasswd usermod -aG sudo test
- 启动实例,等待2分钟后尝试用test用户SSH登录
- 登录成功后排查原用户的
~/.ssh/authorized_keys权限、sshd_config历史修改记录,确认系统升级过程中是否新增了登录IP限制、用户白名单规则
若启动脚本无效,可通过挂载磁盘修复
- 停止故障实例,卸载其系统盘
- 创建一个临时Debian实例,将故障系统盘作为数据盘挂载到临时实例
- 登录临时实例,挂载故障磁盘的根分区,依次检查以下内容:
/var/log/auth.log:查看SSH连接失败的具体报错,确认是密钥认证失败还是服务主动拒绝连接/etc/ssh/sshd_config:检查是否存在AllowUsers/DenyUsers配置限制了登录用户,是否开启了公钥认证- 对应用户主目录下的
.ssh/authorized_keys:确认公钥内容正确,文件权限为600,.ssh目录权限为700
- 修复完成后卸载磁盘,挂载回原实例启动即可
内容的提问来源于stack exchange,提问作者Ankur Raiyani
相关产品推荐
相关产品推荐

