Prometheus向vminsert远程写入时遇Context deadline exceeded错误
问题:Prometheus远程写入vminsert时出现Context Deadline Exceeded
错误日志
ts=2024-09-10T12:10:17.827Z caller=dedupe.go:112 component=remote level=info remote_name=409e40 url=http://x.x.x.x:8480/insert/0/prometheus/api/v1/write msg="Remote storage resharding" from=272 to=500 ts=2024-09-10T12:10:59.892Z caller=dedupe.go:112 component=remote level=warn remote_name=409e40 url=http://x.x.x.x:8480/insert/0/prometheus/api/v1/write msg="Failed to send batch, retrying" err="Post \"http://x.x.x.x:8480/insert/0/prometheus/api/v1/write\": context deadline exceeded"
注意:问题出在远程写入环节,而非指标采集环节,当前采集100余台服务器指标。
现有配置
Prometheus remote_write配置
remote_write: - url: "http://x.x.x.x:8480/insert/0/prometheus/api/v1/write" queue_config: max_shards: 500 min_shards: 8 tls_config: insecure_skip_verify: true
Prometheus Deployment资源与参数
containers: - name: prometheus image: prom/prometheus args: - "--storage.tsdb.retention.time=1h" - "--config.file=/etc/prometheus/prometheus.yml" - "--storage.tsdb.path=/prometheus" - "--storage.tsdb.retention.size=5GB" ports: - containerPort: 9090 resources: requests: cpu: 0.5 memory: 4Gi limits: cpu: 3 memory: 18Gi
vminsert Deployment配置
apiVersion: apps/v1 kind: Deployment metadata: name: vminsert namespace: monitor-system spec: replicas: 3 selector: matchLabels: app: vminsert template: metadata: labels: app: vminsert spec: containers: - name: vminsert image: victoriametrics/vminsert args: - "-maxConcurrentInserts=4096" - "-insert.maxQueueDuration=15m" - "-replicationFactor=2" - -storageNode=vmstorage-0.vmstorage.monitor-system.svc.cluster.local:8400 - -storageNode=vmstorage-1.vmstorage.monitor-system.svc.cluster.local:8400 ports: - containerPort: 8480 name: http-insert
已尝试方案
- 扩容vminsert资源配置,问题未解决
- 将Prometheus远程写入分片数调整至1500,问题仍未解决
解决方案建议
1. 优化Prometheus远程写入的超时与批处理配置
默认远程写入超时时间较短,批处理参数可能不匹配当前数据量,可在remote_write中补充以下配置:
remote_write: - url: "http://x.x.x.x:8480/insert/0/prometheus/api/v1/write" timeout: 30s # 延长单次请求超时时间 queue_config: max_shards: 500 min_shards: 8 max_samples_per_send: 10000 # 调整单批次发送的样本数 batch_send_deadline: 10s # 批次发送的截止时间 min_backoff: 1s # 最小重试间隔 max_backoff: 5s # 最大重试间隔 tls_config: insecure_skip_verify: true
timeout:避免因单批次数据量大导致请求超时max_samples_per_send:平衡单请求大小与请求次数,避免过大或过小batch_send_deadline:防止队列长期堆积未发送的样本
2. 排查vmstorage的性能瓶颈
vminsert仅负责转发数据,最终写入压力落在vmstorage节点,需检查:
- vmstorage的CPU、内存使用率是否超限
- 磁盘IO状态:通过
iostat查看读写延迟、IOPS是否达到瓶颈,磁盘空间是否充足 - vmstorage核心指标:
vm_storage_data_points_inserted_total(写入速率)、vm_storage_disk_write_seconds_total(磁盘写入耗时),判断是否存在写入延迟 - 若磁盘性能不足,可更换SSD或扩容磁盘容量
3. 检查Prometheus与vminsert的网络链路
- 在Prometheus容器内执行
curl -w "%{time_total}\n" http://x.x.x.x:8480/insert/0/prometheus/api/v1/write,测试请求响应时间是否正常 - 验证K8s Service配置:确认vminsert的Service是否正确关联Pod,负载均衡是否均匀分配流量
- 检查网络策略:是否存在限制Prometheus与vminsert通信的规则
4. 调整vminsert的队列与转发参数
当前-insert.maxQueueDuration=15m设置过大,易导致队列积压,可优化参数:
args: - "-maxConcurrentInserts=4096" - "-insert.maxQueueDuration=2m" # 缩短队列等待时间,减少数据积压 - "-insert.maxQueueSize=100000" # 限制队列最大容量,防止内存溢出 - "-maxInsertRequestSize=67108864" # 允许最大64MB的请求体,适配大批次样本 - "-replicationFactor=2" - -storageNode=vmstorage-0.vmstorage.monitor-system.svc.cluster.local:8400 - -storageNode=vmstorage-1.vmstorage.monitor-system.svc.cluster.local:8400
5. 监控Prometheus远程写入指标
通过Prometheus自身指标定位问题:
prometheus_remote_storage_samples_pending_total:查看队列中待发送的样本数,持续增长说明写入能力不足prometheus_remote_storage_succeeded_samples_total:对比总采集样本数,判断写入成功率prometheus_remote_storage_retries_total:频繁重试说明写入链路存在不稳定因素
内容的提问来源于stack exchange,提问作者Akash Prajapati
相关产品推荐
相关产品推荐

