K8s中Redis 6.0.15单机版主从同步触发服务超时故障排查
Kubernetes单机Redis无规律停止响应故障排查
问题背景
- 部署环境:Kubernetes集群内,基于
bitnami/redis:6.0.15镜像部署单机版Redis作为缓存服务 - 自定义启动配置:
MASTER=trueREDIS_AOF_ENABLED=no
- 故障现象:服务无规律停止响应,故障发生后客户端请求队列持续积压,必须删除对应Pod重启Redis才能恢复,否则服务永久不可用
故障现场日志
Redis服务端日志
Jul 5 13:30:27 redis-0 redis 1:M 05 Jul 2022 11:30:27.060 * 10000 changes in 60 seconds. Saving... Jul 5 13:30:27 redis-0 redis 1:M 05 Jul 2022 11:30:27.090 * Background saving started by pid 364 Jul 5 13:31:34 redis-0 redis 364:C 05 Jul 2022 11:31:34.307 * DB saved on disk Jul 5 13:31:34 redis-0 redis 364:C 05 Jul 2022 11:31:34.341 * RDB: 431 MB of memory used by copy-on-write Jul 5 13:31:34 redis-0 redis 1:M 05 Jul 2022 11:31:34.488 * Background saving terminated with success Jul 5 13:32:35 redis-0 redis 1:M 05 Jul 2022 11:32:35.022 * 10000 changes in 60 seconds. Saving... Jul 5 13:32:35 redis-0 redis 1:M 05 Jul 2022 11:32:35.052 * Background saving started by pid 365 ----- Jul 5 13:32:40 redis-0 redis 1:S 05 Jul 2022 11:32:40.436 * Before turning into a replica, using my own master parameters to synthesize a cached master: I may be able to synchronize with the new master with just a partial transfer. Jul 5 13:32:40 redis-0 redis 1:S 05 Jul 2022 11:32:40.436 * REPLICAOF 178.20.40.200:8886 enabled (user request from 'id=71457 addr=10.0.16.46:14072 fd=12 name= age=0 idle=0 flags=N db=0 sub=0 psub=0 multi=-1 qbuf=47 qbuf-free=32721 argv-mem=24 obl=0 oll=0 omem=0 tot-mem=61488 events=r cmd=slaveof user=default') Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.316 * Connecting to MASTER 178.20.40.200:8886 Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.316 * MASTER <-> REPLICA sync started Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.362 * Non blocking connect for SYNC fired the event. Jul 5 13:32:41 redis-0 redis Error 1:S 05 Jul 2022 11:32:41.409 # Error reply to PING from master: '-Reading from master: Connection reset by peer' Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.316 * Connecting to MASTER 178.20.40.200:8886 Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.317 * MASTER <-> REPLICA sync started Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.366 * Non blocking connect for SYNC fired the event. Jul 5 13:32:42 redis-0 redis Error 1:S 05 Jul 2022 11:32:42.415 # Error reply to PING from master: '-Reading from master: Connection reset by peer' Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.317 * Connecting to MASTER 178.20.40.200:8886 Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.317 * MASTER <-> REPLICA sync started Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.366 * Non blocking connect for SYNC fired the event. Jul 5 13:32:43 redis-0 redis Error 1:S 05 Jul 2022 11:32:43.416 # Error reply to PING from master: '-Reading from master: Connection reset by peer' Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.320 * Connecting to MASTER 178.20.40.200:8886 Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.320 * MASTER <-> REPLICA sync started Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.370 * Non blocking connect for SYNC fired the event.
客户端状态采集
next: GET 6126674261995698486, inst: 1, qu: 0, // queue => waiting operations qs: 17, aw: False, rs: ReadAsync, ws: Idle, in: 0, // bytes waiting from input stream in-pipe: 0, out-pipe: 0, serverEndpoint: redis.default.svc.cluster.local:6379, mc: 1/1/0, mgr: 10 of 10 available, // tread pool clientName: production-9bbd94544-nlmv7, IOCP: (Busy=0,Free=1000,Min=5,Max=1000), // no busy threads WORKER: (Busy=14,Free=32753,Min=256,Max=32767), v: 2.2.4.27433
Timeout performing GET (3000ms), next: 2865582319381864083, inst: 0, qu: 0, qs: 333, aw: False, rs: ReadAsync, ws: Idle, in: 0, in-pipe: 0, out-pipe: 0, serverEndpoint: redis.default.svc.cluster.local:6379, mc: 1/1/0, mgr: 10 of 10 available, clientName: production-58c7874fd8-tdcpz, IOCP: (Busy=0,Free=1000,Min=1,Max=1000), WORKER: (Busy=3,Free=32764,Min=256,Max=32767), v: 2.2.4.27433
next: GET 6126674261995698486, inst: 47, qu: 0, qs: 21368, aw: False, rs: ReadAsync, ws: Idle, in: 0, in-pipe: 0, out-pipe: 0, serverEndpoint: redis.default.svc.cluster.local:6379, mc: 1/1/0, mgr: 10 of 10 available, clientName: production-9bbd94544-nlmv7, IOCP: (Busy=0,Free=1000,Min=5,Max=1000), WORKER: (Busy=162,Free=32605,Min=256,Max=32767), v: 2.2.4.27433
根因定位
这是典型的Redis未授权访问被攻击场景,和RDB后台保存没有直接关系:
- 日志明确显示,RDB保存完成5秒后,IP为
10.0.16.46的未授权连接向Redis发送了SLAVEOF 178.20.40.200:8886命令,将原本的主节点Redis强制切换为外部恶意IP的从节点。 - Redis切换为从节点后,会停止响应正常业务读写请求,清空现有数据,持续尝试连接配置的主节点做全量同步。日志里反复出现的连接重置错误,是因为攻击者控制的
178.20.40.200:8886本身就是扫描用的恶意节点,不会正常响应主从同步请求,导致Redis一直卡在从节点同步重试状态,完全无法对外提供服务。 - 客户端状态里
in=0、qs队列持续积压的表现,完全匹配Redis卡在同步状态时不读取、不响应客户端请求的特征。 - 重启Pod能临时恢复,是因为重启时启动参数指定了
MASTER=true,Redis会重新以主节点模式启动;但如果不修复安全漏洞,攻击者会再次扫描到端口,重复执行恶意命令打挂服务。
排查步骤
- 故障复现时直接连接Redis执行
INFO replication,确认实例角色是否为slave,核对master_host、master_port是否匹配日志中的恶意地址,即可验证问题。 - 检查Redis配置,确认是否未配置
requirepass密码,或密码为弱密码。 - 检查K8s Service、Ingress配置,确认Redis 6379端口是否被错误暴露到公网,或未做访问控制允许集群内任意IP访问。
- 检查Redis配置,确认是否未对
SLAVEOF、CONFIG、FLUSHALL等高危命令做禁用/重命名处理。
修复方案
- 立即为Redis设置强复杂度的访问密码,禁止使用空密码、弱密码。
- 配置Kubernetes NetworkPolicy,仅允许业务所在命名空间的Pod访问Redis 6379端口,禁止端口对外网暴露,限制集群内非授权IP的访问权限。
- 在Redis配置中禁用/重命名高危命令,参考配置如下:
rename-command SLAVEOF "" rename-command CONFIG "随机生成的复杂字符串别名" rename-command FLUSHALL "" rename-command FLUSHDB "" rename-command DEBUG "" - 单机部署场景下直接在启动配置中添加
replicaof no one锁定实例角色,禁止运行时被修改为从节点。 - 故障发生时无需重启Pod,连接Redis执行
SLAVEOF NO ONE即可将实例切回主节点模式,配合本地持久化的RDB文件可快速恢复业务,避免重启导致的数据丢失风险。 - 升级
bitnami/redis镜像到最新稳定补丁版本,修复已知的安全漏洞。
内容的提问来源于stack exchange,提问作者Sven Boris Bornemann
相关产品推荐
相关产品推荐

