Prometheus+Docker服务发现:可迁移微服务监控方案咨询
Great question—you’re right to revisit this since 2017, because the ecosystem has evolved a lot to address exactly the cross-network migration and service discovery pain points you’re describing. Let’s walk through the options:
First: Why DNS A Records Fall Short for Cross-Docker Network Scenarios
You’re correct that DNS A records were a go-to in 2017, but they struggle with Docker’s network isolation. When a microservice moves to a different Docker network, the DNS entry might not propagate correctly across networks (especially if using default bridge networks, which are isolated per container), or Prometheus might cache old DNS records, leading to failed scrapes. This makes DNS A records a brittle choice for dynamic, migratable services.
Modern Solutions to the Problem
Thankfully, there are several robust, automated options now depending on your infrastructure:
1. Docker Swarm Service Discovery (If You Use Swarm)
If you’re orchestrating services with Docker Swarm, Prometheus has built-in docker_sd_configs that automatically discover Swarm services and their instances—regardless of which network they’re migrated to. It pulls directly from the Docker API to get up-to-date service endpoints.
Example configuration:
scrape_configs: - job_name: 'docker-swarm-services' docker_sd_configs: - host: unix:///var/run/docker.sock # Use TCP if Prometheus is on a different node role: service relabel_configs: - source_labels: [__meta_docker_service_name] target_label: job
2. Kubernetes Service Discovery (If You Use K8s)
For Kubernetes environments, kubernetes_sd_configs is the gold standard. Prometheus automatically detects pods, services, and endpoints as they’re scheduled or migrated across nodes/networks. No manual updates needed—everything is handled via the Kubernetes API.
3. Service Registries (e.g., Consul, etcd)
If you’re not using an orchestrator, a service registry like Consul solves this problem elegantly. Your microservices register themselves with Consul when they start (and deregister when they stop), and Prometheus uses consul_sd_configs to pull the latest list of targets. This works across any network as long as Prometheus can reach the Consul server.
Example configuration:
scrape_configs: - job_name: 'consul-discovered-services' consul_sd_configs: - server: 'consul-server:8500' services: ['payment-service', 'user-service'] # Target specific services, or leave empty for all
Is file_sd_config the Best Fallback?
If none of the above options fit your setup (e.g., you’re running services without orchestration or a registry), then yes, file_sd_config is an excellent, reliable choice. It’s simple to implement, flexible, and doesn’t require any additional infrastructure.
Here’s how it works:
- You maintain a JSON/YAML file listing your service targets.
- Prometheus periodically reads this file (default every 5 minutes, adjustable with
refresh_interval) and updates its scrape targets automatically—no restart required.
Example configuration:
scrape_configs: - job_name: 'file-based-microservices' file_sd_configs: - files: - '/etc/prometheus/targets/microservices.json' refresh_interval: 1m # Refresh every minute for faster updates
Sample targets file (microservices.json):
[ { "targets": ["payment-service:9090", "user-service:9090"], "labels": { "environment": "production" } } ]
You can automate updates to this file using scripts, configuration management tools (Ansible, SaltStack), or even custom code that triggers when a service migrates. It’s a low-overhead way to keep Prometheus in sync with your dynamic services.
Final Recommendation
Prioritize using orchestrator-native service discovery (Swarm/K8s) or a service registry if possible—these are the most automated and least error-prone options. If those aren’t feasible, file_sd_config is absolutely your best bet for maintaining control over scrape targets while supporting service migration/replication.
内容的提问来源于stack exchange,提问作者MLGonc

