You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Prometheus中不同Job或Target设置差异化数据保留时长以缩减数据库体积

Hey there! Great question—differentiating retention policies per job or target in Prometheus is a super common need when you want to trim storage costs without ditching critical metrics. Let’s break down the most practical, battle-tested ways to make this happen:

1. Multiple Prometheus Instances (The Straightforward Approach)

Prometheus doesn’t natively support per-job retention out of the box, so running separate instances is the simplest way to enforce different rules. Each instance will scrape only a subset of your jobs/targets and use its own retention setting.

Here’s how to set it up:

  • Instance 1 (Critical Workloads): Configure it to scrape high-priority targets (like your core API, database, or auth service) with a long retention window. Start it with:
    prometheus --config.file=prometheus-core.yml --storage.tsdb.retention.time=1y
    
    In prometheus-core.yml, your scrape_configs should only include the critical jobs—use explicit job names or label filters to narrow things down:
    scrape_configs:
      - job_name: "core-api"
        static_configs:
          - targets: ["api-gateway:9090"]
      - job_name: "database"
        static_configs:
          - targets: ["postgres-exporter:9187"]
    
  • Instance 2 (Low-Priority/Workloads): For temporary test services or non-critical monitoring, use a short retention window:
    prometheus --config.file=prometheus-non-core.yml --storage.tsdb.retention.time=7d
    
    Its scrape_configs will target only the low-priority jobs.

Pro Tip: Use Prometheus Federation to aggregate data from both instances into a single query endpoint if you need unified dashboards. You can also point your Alertmanager at both instances to cover all alerts.

2. Remote Write + Tiered Storage (The Flexible Approach)

If you don’t want to manage multiple Prometheus instances, leverage Prometheus’s remote_write feature to route metrics to different storage backends, each with its own retention policy.

For example:

  • Send core job metrics to a long-term storage (like Thanos Store or M3DB) with 1-year retention.
  • Send low-priority metrics to a cheap, short-term storage (like a lightweight Prometheus instance or local TSDB) with 7-day retention.

Here’s a sample prometheus.yml config to implement this:

remote_write:
  # Route core jobs to long-term storage
  - url: "http://long-term-storage:9090/api/v1/write"
    write_relabel_configs:
      - source_labels: [job]
        regex: "core-api|database|auth-service"
        action: keep

  # Route non-core jobs to short-term storage
  - url: "http://short-term-storage:9090/api/v1/write"
    write_relabel_configs:
      - source_labels: [job]
        regex: "test-service|temp-monitor|dev-envs"
        action: keep

Each storage backend handles its own retention: Thanos can be configured with --retention.resolution-raw=365d, while the short-term Prometheus instance uses --storage.tsdb.retention.time=7d.

3. Prometheus Operator (Kubernetes Environments)

If you’re running on Kubernetes with the Prometheus Operator, you can use custom resources to enforce per-job/target retention policies cleanly.

  • Create two separate Prometheus CRs, each with its own retention window and a selector to match specific PodMonitor/ServiceMonitor resources:
    # Core Prometheus instance (1-year retention)
    apiVersion: monitoring.coreos.com/v1
    kind: Prometheus
    metadata:
      name: prometheus-core
    spec:
      retention: 365d
      scrapeSelector:
        matchLabels:
          monitoring-tier: core
      serviceAccountName: prometheus
      # Other configs like storage volumes, replicas, etc.
    
    ---
    # Non-core Prometheus instance (7-day retention)
    apiVersion: monitoring.coreos.com/v1
    kind: Prometheus
    metadata:
      name: prometheus-non-core
    spec:
      retention: 7d
      scrapeSelector:
        matchLabels:
          monitoring-tier: non-core
      serviceAccountName: prometheus
    
  • Then tag your PodMonitor/ServiceMonitor resources with the corresponding monitoring-tier label:
    apiVersion: monitoring.coreos.com/v1
    kind: PodMonitor
    metadata:
      name: core-api-monitor
      labels:
        monitoring-tier: core
    spec:
      selector:
        matchLabels:
          app: core-api
      podMetricsEndpoints:
        - port: metrics
    
4. Bonus: Aggregate Low-Priority Metrics

If you want to stick with a single Prometheus instance, you can reduce storage usage for low-priority jobs by aggregating their metrics into coarser intervals using recording rules, then drop the raw high-frequency data.

Example recording rule:

groups:
  - name: non-core-metrics-aggregation
    rules:
      - record: non_core:cpu_usage:avg_5m
        expr: avg_over_time(non_core_cpu_usage[5m])
        labels:
          aggregated: "true"

Then use metric_relabel_configs in your scrape config to drop the raw metrics:

scrape_configs:
  - job_name: "test-service"
    static_configs:
      - targets: ["test-service:9090"]
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: "non_core_cpu_usage"
        action: drop

This way, you only store the aggregated 5-minute metrics, which take up far less space even with the same global retention window.


A few final notes to keep in mind:

  • Always monitor your storage usage after implementing these changes to adjust retention windows as needed.
  • For remote write setups, ensure you have redundancy (like replication) to avoid data loss if a storage backend goes down.
  • Multiple instances can add operational overhead, so weigh that against the storage savings you’ll get.

内容的提问来源于stack exchange,提问作者byteUI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 19:42:47