SolrCloud 9.1 长期运行后查询性能下降问题排查求助
SolrCloud 9.1查询性能随时间下降问题分析
集群环境
- 基于3台r5.xlarge EC2实例部署SolrCloud 9.1集群
- 采用EFS共享磁盘+stunnel架构,Solr数据目录及ZooKeeper均部署在共享系统中
- 每日新增约20000份文档,集群需同时处理索引操作、常规查询及新功能特殊查询
性能异常现象
刚重启Solr时,目标特殊查询执行速度极快,但运行时间越久,查询速度逐渐变慢:
重启后(time1)请求耗时仅2.162秒,经过一夜运行后(time2)响应耗时达到18.137秒
典型性能对比数据:
businessId, all count, reduced count, time1, time2 7016274253,8433,4769,2.162,18.137
问题查询参数
该慢查询的请求详情如下:
url: "http://xxx.aws01.hibu.int:8983/solr/calls/select", params: { q: `business_id:${businessId} AND call_day:[20230101 TO 20240101}`, fl: "business_id, call_id, call_day, call_date, dialog_merged, call_callerno, call_duration, call_status, caller_name, caller_address, caller_state, caller_city, caller_zip", rows: limit, start: 0, group: true, "group.main": true, "group.field": "call_callerno", sort: "call_day desc" }
当前Commit配置
客户端仅使用softCommit,具体配置:
<autoCommit> <maxTime>180000</maxTime> <maxSize>512m</maxSize> <openSearcher>false</openSearcher> </autoCommit> <autoSoftCommit> <maxTime>10000</maxTime> </autoSoftCommit>
Metrics观测数据
通过/solr/admin/metrics监控发现,不同分片的查询耗时差异显著:
分片1请求耗时统计
"QUERY./select.requestTimes":{ "count":4577, "meanRate":0.09252592498547889, "1minRate":0.07171534322545538, "5minRate":0.056511876693544336, "15minRate":0.05780642380709814, "min_ms":5.607831, "max_ms":35447.542165, "mean_ms":12.160278707076563, "median_ms":5.988622, "stddev_ms":14.871542074236968, "p75_ms":6.307839, "p95_ms":42.103719, "p99_ms":42.103719, "p999_ms":98.124416},
分片2请求耗时统计(异常偏高)
"QUERY./select.requestTimes":{ "count":4486, "meanRate":0.09405828676729713, "1minRate":0.09345322035169516, "5minRate":0.062102810330670666, "15minRate":0.05520043855292057, "min_ms":5.666243, "max_ms":34713.632736, "mean_ms":272.95919728573585, "median_ms":6.101441, "stddev_ms":813.2470530531275, "p75_ms":7.397941, "p95_ms":3392.606168, "p99_ms":3392.606168, "p999_ms":3392.606168},
核心疑问
- 查询性能随时间推移持续下降的原因是什么?
- 重启后性能恢复的根本逻辑是什么?
- 怀疑问题出在索引或缓存配置上,求验证方向及优化建议
内容的提问来源于stack exchange,提问作者laloumen
相关产品推荐
相关产品推荐

