Varnish内存持续上涨致EC2实例崩溃,求调优解决方案
Varnish内存持续增长导致EC2实例崩溃的排查与解决
问题背景
我们在两台**t4g.medium(4GB内存)**的EC2实例上运行Varnish,遇到内存占用持续增长直至实例崩溃的问题。已尝试的优化措施:
- 配置
malloc分配2048MB缓存空间,但未影响varnishd内存占用 - 将最小线程数从200降至100,无明显改善
当前Varnish系统服务配置
[Unit] Description=Varnish Cache, a high-performance HTTP accelerator After=network-online.target [Service] Type=simple Environment="MALLOC_CONF=thp:never,narenas:2" # Maximum number of open files (for ulimit -n) LimitNOFILE=131072 # Locked shared memory - should suffice to lock the shared memory log # (varnishd -l argument) # Default log size is 80MB vsl + 1M vsm + header -> 82MB # unit is bytes LimitMEMLOCK=85983232 # Enable this to avoid "fork failed" on reload. #TasksMax=infinity # Maximum size of the corefile. #LimitCORE=infinity ExecStart=/usr/sbin/varnishd -j unix,user=vcache -F -a :80 -T :6082 -f /etc/varnish/default.vcl -S /etc/varnish/secret -p vcc_allow_inline_c=on -p feature=+esi_ignore_other_elements -p feature=+esi_disable_xml_check -p http_max_hdr=128 -p http_resp_hdr_len=42000 -p http_resp_size=74768 -p workspace_client=256k -p workspace_backend=256k -p feature=+esi_ignore_https -p thread_pool_min=50 -s malloc,2048m ExecReload=/usr/sbin/varnishreload ProtectSystem=full ProtectHome=true PrivateTmp=true PrivateDevices=true [Install] WantedBy=multi-user.target
Varnish状态统计(varnishstat输出)
MGT.uptime 0+03:50:37 MAIN.uptime 0+03:50:38 MAIN.sess_conn 11765 MAIN.client_req 93915 MAIN.cache_hit 11321 MAIN.cache_hitmiss 9224 MAIN.cache_miss 44954 MAIN.backend_conn 5675 MAIN.backend_reuse 74129 MAIN.backend_recycle 79616 MAIN.fetch_head 118 MAIN.fetch_length 10725 MAIN.fetch_chunked 50532 MAIN.fetch_none 5310 MAIN.fetch_304 13083 MAIN.fetch_failed 1 MAIN.pools 2 MAIN.threads 100 MAIN.threads_created 100 MAIN.busy_sleep 305 MAIN.busy_wakeup 305 MAIN.n_object 23839 MAIN.n_objectcore 23863 MAIN.n_objecthead 20855 MAIN.n_backend 3 MAIN.n_lru_nuked 602 MAIN.s_sess 11765 MAIN.s_pipe 33 MAIN.s_pass 34815 MAIN.s_fetch 79769 MAIN.s_synth 2856 MAIN.s_req_hdrbytes 219.91M MAIN.s_req_bodybytes 9.47M MAIN.s_resp_hdrbytes 49.07M MAIN.s_resp_bodybytes 8.14G MAIN.s_pipe_hdrbytes 24.02K MAIN.s_pipe_out 4.25M MAIN.sess_closed 1564 MAIN.sess_closed_err 9501 MAIN.backend_req 82081 MAIN.n_vcl 1 MAIN.bans 1 MAIN.vmods 2 MAIN.n_gzip 33598 MAIN.n_gunzip 29012 SMA.s0.c_req 352442 SMA.s0.c_fail 917 SMA.s0.c_bytes 4.79G SMA.s0.c_freed 2.79G SMA.s0.g_alloc 147586 SMA.s0.g_bytes 2.00G SMA.s0.g_space 124.88K SMA.Transient.c_req 262975 SMA.Transient.c_bytes 3.03G SMA.Transient.c_freed 3.02G SMA.Transient.g_alloc 13436 SMA.Transient.g_bytes 11.20M VBE.boot.web_asg_10_0_2_23.happy ffffffffff VBE.boot.web_asg_10_0_2_23.bereq_hdrbytes 64.35M VBE.boot.web_asg_10_0_2_23.bereq_bodybytes 1.38M VBE.boot.web_asg_10_0_2_23.beresp_hdrbytes 23.24M VBE.boot.web_asg_10_0_2_23.beresp_bodybytes 1.08G VBE.boot.web_asg_10_0_2_23.pipe_hdrbytes 9.65K VBE.boot.web_asg_10_0_2_23.pipe_in 1.61M VBE.boot.web_asg_10_0_2_23.conn 2 VBE.boot.web_asg_10_0_2_23.req 27608 VBE.boot.web_asg_10_0_1_174.happy ffffffffff VBE.boot.web_asg_10_0_1_174.bereq_hdrbytes 65.10M VBE.boot.web_asg_10_0_1_174.bereq_bodybytes 6.66M VBE.boot.web_asg_10_0_1_174.beresp_hdrbytes 23.25M VBE.boot.web_asg_10_0_1_174.beresp_bodybytes 1.13G VBE.boot.web_asg_10_0_1_174.pipe_hdrbytes 5.54K VBE.boot.web_asg_10_0_1_174.pipe_in 973.57K VBE.boot.web_asg_10_0_1_174.conn 4 VBE.boot.web_asg_10_0_1_174.req 27608 VBE.boot.web_asg_10_0_3_248.happy ffffffffff VBE.boot.web_asg_10_0_3_248.bereq_hdrbytes 64.92M VBE.boot.web_asg_10_0_3_248.bereq_bodybytes 1.47M VBE.boot.web_asg_10_0_3_248.beresp_hdrbytes 23.37M VBE.boot.web_asg_10_0_3_248.beresp_bodybytes 1.12G VBE.boot.web_asg_10_0_3_248.pipe_hdrbytes 10.33K VBE.boot.web_asg_10_0_3_248.pipe_in 1.68M VBE.boot.web_asg_10_0_3_248.conn 3 VBE.boot.web_asg_10_0_3_248.req 27609
解决方案建议
1. 优化缓存回收机制
从统计数据看,SMA.s0.g_bytes已达2GB缓存上限,但MAIN.n_lru_nuked仅602,说明LRU回收未充分触发:
- 检查VCL中
beresp.ttl配置,缩短非热点内容的TTL,避免大量旧对象堆积 - 调整
lru_interval参数(默认2秒),提高LRU扫描频率,加快旧对象回收
2. 降低工作区内存开销
当前workspace_client=256k和workspace_backend=256k设置过高,每个线程都会占用对应内存:
- 将
workspace_client降至64k,workspace_backend降至128k,减少单线程内存占用 - 添加
thread_pool_max=200参数,限制线程总数,避免内存过度消耗
3. 减少不必要的Pass请求
MAIN.s_pass达34815次,占总请求近37%,大量Pass请求会增加临时内存负载:
- 排查VCL规则,优化缓存逻辑,将可缓存的请求纳入缓存范围
- 对必须Pass的请求,确保临时存储(Transient SMA)配置合理,避免内存溢出
4. 调整系统内存限制
当前LimitMEMLOCK仅覆盖日志内存,需让Varnish能锁定缓存内存:
- 将
LimitMEMLOCK改为2147483648(2GB),与缓存分配大小一致 - 取消
TasksMax的注释,设置为TasksMax=infinity,避免重载时fork失败
5. 排查内存泄漏
若上述优化后仍有内存增长,需进一步排查:
- 用
varnishlog跟踪大对象请求,查看是否有异常对象无法被回收 - 升级Varnish到最新稳定版本,修复已知的内存泄漏bug
内容的提问来源于stack exchange,提问作者Andrea
相关产品推荐
相关产品推荐

