You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K8s中Redis 6.0.15单机版主从同步触发服务超时故障排查

Kubernetes单机Redis无规律停止响应故障排查

问题背景

  • 部署环境:Kubernetes集群内,基于bitnami/redis:6.0.15镜像部署单机版Redis作为缓存服务
  • 自定义启动配置:
    • MASTER=true
    • REDIS_AOF_ENABLED=no
  • 故障现象:服务无规律停止响应,故障发生后客户端请求队列持续积压,必须删除对应Pod重启Redis才能恢复,否则服务永久不可用

故障现场日志

Redis服务端日志

Jul 5 13:30:27 redis-0 redis 1:M 05 Jul 2022 11:30:27.060 * 10000 changes in 60 seconds. Saving...
Jul 5 13:30:27 redis-0 redis 1:M 05 Jul 2022 11:30:27.090 * Background saving started by pid 364
Jul 5 13:31:34 redis-0 redis 364:C 05 Jul 2022 11:31:34.307 * DB saved on disk
Jul 5 13:31:34 redis-0 redis 364:C 05 Jul 2022 11:31:34.341 * RDB: 431 MB of memory used by copy-on-write
Jul 5 13:31:34 redis-0 redis 1:M 05 Jul 2022 11:31:34.488 * Background saving terminated with success
Jul 5 13:32:35 redis-0 redis 1:M 05 Jul 2022 11:32:35.022 * 10000 changes in 60 seconds. Saving...
Jul 5 13:32:35 redis-0 redis 1:M 05 Jul 2022 11:32:35.052 * Background saving started by pid 365
-----
Jul 5 13:32:40 redis-0 redis 1:S 05 Jul 2022 11:32:40.436 * Before turning into a replica, using my own master parameters to synthesize a cached master: I may be able to synchronize with the new master with just a partial transfer.
Jul 5 13:32:40 redis-0 redis 1:S 05 Jul 2022 11:32:40.436 * REPLICAOF 178.20.40.200:8886 enabled (user request from 'id=71457 addr=10.0.16.46:14072 fd=12 name= age=0 idle=0 flags=N db=0 sub=0 psub=0 multi=-1 qbuf=47 qbuf-free=32721 argv-mem=24 obl=0 oll=0 omem=0 tot-mem=61488 events=r cmd=slaveof user=default')
Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.316 * Connecting to MASTER 178.20.40.200:8886
Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.316 * MASTER <-> REPLICA sync started
Jul 5 13:32:41 redis-0 redis 1:S 05 Jul 2022 11:32:41.362 * Non blocking connect for SYNC fired the event.
Jul 5 13:32:41 redis-0 redis Error 1:S 05 Jul 2022 11:32:41.409 # Error reply to PING from master: '-Reading from master: Connection reset by peer'
Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.316 * Connecting to MASTER 178.20.40.200:8886
Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.317 * MASTER <-> REPLICA sync started
Jul 5 13:32:42 redis-0 redis 1:S 05 Jul 2022 11:32:42.366 * Non blocking connect for SYNC fired the event.
Jul 5 13:32:42 redis-0 redis Error 1:S 05 Jul 2022 11:32:42.415 # Error reply to PING from master: '-Reading from master: Connection reset by peer'
Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.317 * Connecting to MASTER 178.20.40.200:8886
Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.317 * MASTER <-> REPLICA sync started
Jul 5 13:32:43 redis-0 redis 1:S 05 Jul 2022 11:32:43.366 * Non blocking connect for SYNC fired the event.
Jul 5 13:32:43 redis-0 redis Error 1:S 05 Jul 2022 11:32:43.416 # Error reply to PING from master: '-Reading from master: Connection reset by peer'
Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.320 * Connecting to MASTER 178.20.40.200:8886
Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.320 * MASTER <-> REPLICA sync started
Jul 5 13:32:44 redis-0 redis 1:S 05 Jul 2022 11:32:44.370 * Non blocking connect for SYNC fired the event.

客户端状态采集

next: GET 6126674261995698486, 
inst: 1, 
qu: 0,  // queue => waiting operations
qs: 17, 
aw: False, 
rs: ReadAsync, 
ws: Idle, 
in: 0, // bytes waiting from input stream
in-pipe: 0, 
out-pipe: 0, 
serverEndpoint: redis.default.svc.cluster.local:6379, 
mc: 1/1/0, 
mgr: 10 of 10 available, // tread pool
clientName: production-9bbd94544-nlmv7, 
IOCP: (Busy=0,Free=1000,Min=5,Max=1000),  // no busy threads
WORKER: (Busy=14,Free=32753,Min=256,Max=32767), 
v: 2.2.4.27433
Timeout performing GET (3000ms), 
next: 2865582319381864083, 
inst: 0, 
qu: 0, 
qs: 333, 
aw: False, 
rs: ReadAsync, 
ws: Idle, 
in: 0, 
in-pipe: 0, 
out-pipe: 0, 
serverEndpoint: redis.default.svc.cluster.local:6379, 
mc: 1/1/0, 
mgr: 10 of 10 available, 
clientName: production-58c7874fd8-tdcpz, 
IOCP: (Busy=0,Free=1000,Min=1,Max=1000), 
WORKER: (Busy=3,Free=32764,Min=256,Max=32767), 
v: 2.2.4.27433 
next: GET 6126674261995698486, 
inst: 47, 
qu: 0, 
qs: 21368, 
aw: False, 
rs: ReadAsync,
 ws: Idle, 
 in: 0, 
 in-pipe: 0, 
 out-pipe: 0, 
 serverEndpoint: redis.default.svc.cluster.local:6379, 
 mc: 1/1/0, 
 mgr: 10 of 10 available, 
 clientName: production-9bbd94544-nlmv7, 
 IOCP: (Busy=0,Free=1000,Min=5,Max=1000), 
 WORKER: (Busy=162,Free=32605,Min=256,Max=32767), 
v: 2.2.4.27433

根因定位

这是典型的Redis未授权访问被攻击场景,和RDB后台保存没有直接关系:

  1. 日志明确显示,RDB保存完成5秒后,IP为10.0.16.46的未授权连接向Redis发送了SLAVEOF 178.20.40.200:8886命令,将原本的主节点Redis强制切换为外部恶意IP的从节点。
  2. Redis切换为从节点后,会停止响应正常业务读写请求,清空现有数据,持续尝试连接配置的主节点做全量同步。日志里反复出现的连接重置错误,是因为攻击者控制的178.20.40.200:8886本身就是扫描用的恶意节点,不会正常响应主从同步请求,导致Redis一直卡在从节点同步重试状态,完全无法对外提供服务。
  3. 客户端状态里in=0、qs队列持续积压的表现,完全匹配Redis卡在同步状态时不读取、不响应客户端请求的特征。
  4. 重启Pod能临时恢复,是因为重启时启动参数指定了MASTER=true,Redis会重新以主节点模式启动;但如果不修复安全漏洞,攻击者会再次扫描到端口,重复执行恶意命令打挂服务。

排查步骤

  • 故障复现时直接连接Redis执行INFO replication,确认实例角色是否为slave,核对master_host、master_port是否匹配日志中的恶意地址,即可验证问题。
  • 检查Redis配置,确认是否未配置requirepass密码,或密码为弱密码。
  • 检查K8s Service、Ingress配置,确认Redis 6379端口是否被错误暴露到公网,或未做访问控制允许集群内任意IP访问。
  • 检查Redis配置,确认是否未对SLAVEOF、CONFIG、FLUSHALL等高危命令做禁用/重命名处理。

修复方案

  • 立即为Redis设置强复杂度的访问密码,禁止使用空密码、弱密码。
  • 配置Kubernetes NetworkPolicy,仅允许业务所在命名空间的Pod访问Redis 6379端口,禁止端口对外网暴露,限制集群内非授权IP的访问权限。
  • 在Redis配置中禁用/重命名高危命令,参考配置如下:
    rename-command SLAVEOF ""
    rename-command CONFIG "随机生成的复杂字符串别名"
    rename-command FLUSHALL ""
    rename-command FLUSHDB ""
    rename-command DEBUG ""
    
  • 单机部署场景下直接在启动配置中添加replicaof no one锁定实例角色,禁止运行时被修改为从节点。
  • 故障发生时无需重启Pod,连接Redis执行SLAVEOF NO ONE即可将实例切回主节点模式,配合本地持久化的RDB文件可快速恢复业务,避免重启导致的数据丢失风险。
  • 升级bitnami/redis镜像到最新稳定补丁版本,修复已知的安全漏洞。

内容的提问来源于stack exchange,提问作者Sven Boris Bornemann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 17:18:24