VictoriaMetrics集群模式下vminsert导入数据返回204但未写入问题
问题背景
已在Kubernetes中搭建VictoriaMetrics集群,vminsert、vmselect、vmstorage的Pod及Service均处于正常运行状态。调用vminsert的导入API返回204 No Content状态码,但数据无法被查询,疑似未成功写入。
集群状态信息
Pod状态
NAME READY STATUS RESTARTS AGE pod/vminsert-79955fd456-f6p5f 1/1 Running 0 2m15s pod/vminsert-79955fd456-qgrlv 1/1 Running 0 2m12s pod/vminsert-79955fd456-skc2x 1/1 Running 0 2m12s pod/vmselect-698556d84c-7z9p7 1/1 Running 0 150m pod/vmselect-698556d84c-8spgr 1/1 Running 0 150m pod/vmselect-698556d84c-sp8n9 1/1 Running 0 150m pod/vmstorage-0 1/1 Running 0 151m pod/vmstorage-1 1/1 Running 0 150m pod/vmstorage-2 1/1 Running 0 150m pod/vmstorage-3 1/1 Running 0 150m pod/vmstorage-4 1/1 Running 0 149m
Service状态
NAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGE service/vminsert LoadBalancer 172.0.0.81 x.x.x.x 8480:30088/TCP,4242:32436/TCP 2m15s service/vmselect LoadBalancer 172.0.0.79 y.y.y.y 8481:32473/TCP 150m service/vmstorage ClusterIP None <none> 8482/TCP,8401/TCP,8400/TCP 151m
API调用详情
调用命令及返回结果:
$ curl -d 'test{foo="bar"} 123' -X POST http://x.x.x.x:8480/insert/0/prometheus/api/v1/import/prometheus -u admin -v Enter host password for user 'admin': * About to connect() to x.x.x.x port 8480 (#0) * Trying x.x.x.x... * Connected to x.x.x.x (x.x.x.x) port 8480 (#0) * Server auth using Basic with user 'admin' > POST /insert/0/prometheus/api/v1/import/prometheus HTTP/1.1 > Authorization: Basic YWRtaW46YWRtaW4= > User-Agent: curl/7.29.0 > Host: x.x.x.x:8480 > Accept: */* > Content-Length: 19 > Content-Type: application/x-www-form-urlencoded > * upload completely sent off: 19 out of 19 bytes < HTTP/1.1 204 No Content < X-Server-Hostname: vminsert-79955fd456-skc2x < Date: Wed, 22 Feb 2023 07:55:23 GMT < * Connection #0 to host x.x.x.x left intact
问题排查步骤
检查vminsert与vmstorage连通性
进入任意vminsert Pod内部,执行curl vmstorage-0.vmstorage:8482/health,验证能否访问vmstorage的健康检查接口。若无法访问,说明DNS解析或网络策略存在问题,导致vminsert无法转发数据到vmstorage。同时查看vminsert日志:kubectl logs <vminsert-pod-name>,排查是否有cannot send data to vmstorage类报错。验证数据格式与时间戳
VictoriaMetrics默认接受Unix毫秒级时间戳,若未指定则使用服务器当前时间。若本地时间与集群时间偏差过大,会导致数据无法查询。可在写入时添加时间戳,比如test{foo="bar"} 123 1677023723000(替换为当前Unix毫秒时间戳)再尝试查询。同时确认数据格式符合Prometheus规范:标签名需匹配[a-zA-Z_][a-zA-Z0-9_]*,值为数字。检查查询方式正确性
使用vmselect的查询接口,访问http://y.y.y.y:8481/select/0/prometheus/api/v1/query?query=test,注意路径中的数据库ID需与写入时的/insert/0/一致。若查询不到,可指定时间范围:http://y.y.y.y:8481/select/0/prometheus/api/v1/query_range?query=test&start=1677019200&end=1677026400&step=15s,避免因时间范围不包含数据导致查询失败。检查vmstorage存储状态
查看vmstorage日志:kubectl logs <vmstorage-pod-name>,排查是否有added 1 samples类写入记录,或存储目录权限不足、磁盘空间不足的错误。进入vmstorage Pod,检查默认存储路径/vmdata下是否有新生成的文件:ls -l /vmdata。验证认证配置一致性
确认vminsert和vmstorage的认证配置是否一致。若vmstorage开启认证,但vminsert未配置对应-storageAuth.username和-storageAuth.password启动参数,会导致数据被静默丢弃,此时vminsert仍可能返回204。
正确的数据写入方法
标准Prometheus格式写入
指定正确的Content-Type为text/plain,示例:curl -d 'test{foo="bar"} 123' -H "Content-Type: text/plain" -X POST http://x.x.x.x:8480/insert/0/prometheus/api/v1/import/prometheus -u admin指定时间戳写入
写入历史数据时添加Unix毫秒时间戳:curl -d 'test{foo="bar"} 123 1677023723000' -H "Content-Type: text/plain" -X POST http://x.x.x.x:8480/insert/0/prometheus/api/v1/import/prometheus -u admin批量写入
一次性写入多条指标,每行一条:curl -d $'test{foo="bar"} 123\ntest{foo="baz"} 456' -H "Content-Type: text/plain" -X POST http://x.x.x.x:8480/insert/0/prometheus/api/v1/import/prometheus -u admin
内容的提问来源于stack exchange,提问作者sjh8763

