Loki作为Grafana数据源加载缓慢问题排查求助
Loki作为Grafana数据源加载缓慢/失败问题排查与解决
问题描述
访问Grafana的Explorer页面时,文件名、jobs、标签等元素加载耗时极久,甚至完全加载失败,"select label"区域加载动画长时间转动。
使用版本
- Loki 2.4.7_4
- Grafana 10.1.1
Loki配置文件
auth_enabled: false server: http_listen_port: 3100 grpc_listen_port: 9096 http_server_read_timeout: 300s http_server_write_timeout: 300s http_server_idle_timeout: 300s common: path_prefix: /tmp/loki storage: filesystem: chunks_directory: /tmp/loki/chunks rules_directory: /tmp/loki/rules replication_factor: 1 ring: kvstore: store: inmemory schema_config: configs: - from: 2020-10-24 store: boltdb-shipper object_store: filesystem schema: v11 index: prefix: index_ period: 24h ruler: alertmanager_url: http://localhost:9093 query_scheduler: max_outstanding_requests_per_tenant: 4096 frontend: max_outstanding_per_tenant: 4096 query_range: parallelise_shardable_queries: true ingester: chunk_idle_period: 15m # This increases the idle period before flushing a chunk max_chunk_age: 2h limits_config: split_queries_by_interval: 5m max_query_parallelism: 64 max_streams_per_user: 30000 max_global_streams_per_user: 0
Loki服务文件
[Unit] Description=Loki service After=network.target [Service] Type=simple User=loki ExecStart=/usr/local/bin/loki-linux-amd64 -config.file /usr/local/bin/config-loki.yml [Install] WantedBy=multi-user.target
优化方案
1. 存储层优化(核心瓶颈)
当前使用boltdb-shipper+本地文件系统作为索引存储,数据量增大后IO性能会成为明显瓶颈:
- 替换索引存储为
tsdb:Loki 2.3+支持的tsdb索引性能远优于boltdb-shipper,修改schema_config:schema_config: configs: - from: 2020-10-24 store: tsdb object_store: filesystem schema: v13 index: prefix: index_ period: 24h - 迁移存储目录:将
/tmp/loki迁移到SSD磁盘挂载目录,避免使用临时目录(/tmp可能被系统定期清理,且性能不稳定),机械硬盘IO性能无法支撑Loki的索引查询需求。
2. 查询并发与传输优化
调整查询调度和前端参数,提升并发处理能力并减少网络开销:
query_scheduler: max_outstanding_requests_per_tenant: 8192 grpc_client_config: max_send_msg_size: 104857600 # 100MB,避免大查询请求被截断 frontend: max_outstanding_per_tenant: 8192 compress_responses: true # 开启响应压缩,降低网络传输耗时 limits_config: split_queries_by_interval: 1m # 缩短查询拆分间隔,减少单任务负载 max_query_parallelism: 128 # 根据CPU核心数调整,建议为核心数的2倍
3. Ingester Chunk参数优化
当前配置会生成过大的chunk,导致查询时需加载更多数据:
ingester: chunk_idle_period: 5m # 缩短chunk空闲刷新周期,减小单个chunk体积 max_chunk_age: 1h # 限制chunk最大存活时间,避免生成超大chunk chunk_target_size: 1536000 # 1.5MB,精准控制chunk大小,提升查询效率
4. 版本升级
Loki 2.4.7属于老旧版本,后续2.8+版本对查询性能、索引存储、稳定性有大量优化,建议升级到最新稳定版。Grafana 10.1.1与新版本Loki兼容性良好,升级后能获得显著性能提升。
5. 系统资源排查
- 用
top/htop监控Loki进程的CPU、内存占用:若CPU持续高负载,说明查询并行度不足或存储IO瓶颈;若内存不足导致频繁swap,需增加内存资源。 - 用
iostat检查磁盘IO负载:若%util接近100%,说明磁盘性能无法满足需求,必须更换SSD或分布式存储。
内容的提问来源于stack exchange,提问作者Saksham Paliwal
相关产品推荐
相关产品推荐

