使用Monstache同步MongoDB至Elasticsearch遇熔断异常求助
问题
环境配置
- 集群配置:2GB内存/30GB存储的Elasticsearch集群
- 初始同步数据量:500万条记录
- 实时同步负载:200请求/分钟
- 搜索查询负载:200请求/分钟
Monstache配置文件
# connection settings # connect to MongoDB using the following URL #mongo-url = "<MONGO_URL>" mongo-url = "<MONGO_URL>/?readConcernLevel=majority&readPreference=secondary" # connect to the Elasticsearch REST API at the following node URLs elasticsearch-urls = ["<ELASTIC_SEARCH_URLS>"] # Array of connection string # frequently required settings # if you need to seed an index from a collection and not just listen and sync changes events # you can copy entire collections or views from MongoDB to Elasticsearch direct-read-namespaces = ["live.chassistypes", "live.chassisowners", "live.containertypes", "live.containerowners", "live.equipment", "live.drivers", "live.customers", "live.loads"] change-stream-namespaces = ["live.chassistypes", "live.chassisowners", "live.containertypes", "live.containerowners", "live.equipment", "live.drivers", "live.customers", "live.loads"] elasticsearch-api-key = "<ELASTIC_SEARCH_API_KEY>" resume = true resume-write-unsafe = false resume-name = "sync-latest" resume-strategy = 1 verbose=true direct-read-split-max = 1000 # Limit the number of documents processed per split direct-read-concur = 2 # Reduce concurrency for direct reads # Server failure handle and auto restart, searcher helth checkup configuration elasticsearch-retry = true # Enable automatic retry on failures elasticsearch-max-seconds = 60 # Maximum time Monstache should wait for a response from Elasticsearch elasticsearch-max-conns = 10 # Maximum number of concurrent connections to Elasticsearch elasticsearch-healthcheck-timeout = 5 # Interval to check if Elasticsearch is healthy # Script middleware for syncing [[pipeline]] namespace = "live.chassistypes" path = "./src/elasticSyncScripts/chassisTypeSync.js" [[pipeline]] namespace = "live.chassisowners" path = "./src/elasticSyncScripts/chassisOwnerSync.js" [[pipeline]] namespace = "live.containertypes" path = "./src/elasticSyncScripts/containerTypesSync.js" [[pipeline]] namespace = "live.containerowners" path = "./src/elasticSyncScripts/containerOwnersSync.js" [[pipeline]] namespace = "live.equipment" path = "./src/elasticSyncScripts/equipmentSync.js" [[pipeline]] namespace = "live.drivers" path = "./src/elasticSyncScripts/driversSync.js" [[pipeline]] namespace = "live.customers" path = "./src/elasticSyncScripts/customersSync.js" [[pipeline]] namespace = "live.loads" path = "./src/elasticSyncScripts/loadSync.js" # Mapping mongo read namespace and elastic store index [[mapping]] namespace="live.chassistypes" index="live-mastersearch-index" [[mapping]] namespace="live.chassisowners" index="live-mastersearch-index" [[mapping]] namespace="live.containertypes" index="live-mastersearch-index" [[mapping]] namespace="live.containerowners" index="live-mastersearch-index" [[mapping]] namespace="live.equipment" index="live-mastersearch-index" [[mapping]] namespace="live.drivers" index="live-mastersearch-index" [[mapping]] namespace="live.customers" index="live-mastersearch-index" [[mapping]] namespace="live.loads" index="live-mastersearch-index" [logs] error="./src/monstacheSyncLogs/error.txt" warn="./src/monstacheSyncLogs/warn.txt" # info="./src/monstacheSyncLogs/info.txt" # trace="./src/monstacheSyncLogs/trace.txt" # stats="./src/monstacheSyncLogs/stats.txt"
错误日志
ERROR 2024/12/20 10:34:51 elastic: bulk processor "monstache" failed but may retry: elastic: Error 429 (Too Many Requests): [parent] Data too large, data for [<http_request>] would be [954104862/909.9mb], which is larger than the limit of [858993459/819.1mb], real usage: [937326936/893.9mb], new bytes reserved: [16777926/16mb], usages [model_inference=0/0b, eql_sequence=0/0b, fielddata=5586/5.4kb, request=0/0b, inflight_requests=38822354/37mb] [type=circuit_breaking_exception] ERROR 2024/12/20 10:35:00 elastic: bulk processor "monstache" failed but may retry: elastic: Error 429 (Too Many Requests): rejected execution of coordinating operation [coordinating_and_primary_bytes=74496147, replica_bytes=25403291, all_bytes=99899438, coordinating_operation_bytes=10284532, max_coordinating_and_primary_bytes=107374182] [type=es_rejected_execution_exception]
分析与解决方案
问题根源
- 内存资源不足:2GB内存的Elasticsearch集群中,父级熔断阈值默认是堆内存的70%(约1.4GB),但实际内存使用已接近900MB,加上新请求需要的内存直接触发熔断。
- 批量写入压力过大:Monstache默认批量配置未做限制,单次批量请求过大,超出Elasticsearch的处理能力。
- 单索引资源竞争:所有MongoDB集合同步到同一个Elasticsearch索引,写入和查询的资源压力高度集中。
具体解决措施
1. 调整Monstache批量写入参数(优先配置)
在Monstache配置文件中添加以下参数,降低批量请求的大小和并发:
# 调整批量写入参数 elasticsearch-bulk-size = 500 # 单次批量请求文档数从默认1000减半 elasticsearch-bulk-concurrency = 2 # 并发批量请求数从默认5降低到2 elasticsearch-bulk-flush-interval = "1s" # 强制1秒刷新一次,避免积累过多待写入文档
2. 优化Elasticsearch内存配置
- 修改
config/jvm.options调整堆内存(2GB机器建议分配1GB堆内存,留1GB给系统和Lucene):-Xms1g -Xmx1g - 调整熔断阈值(谨慎操作,避免OOM),在
config/elasticsearch.yml中添加:indices.breaker.total.limit: 80% # 总熔断阈值从默认70%提高到80% network.breaker.inflight_requests.limit: 50% # inflight请求熔断阈值提高到50%
3. 拆分索引,分散资源压力
将不同业务类型的集合拆分到多个Elasticsearch索引,例如:
- 设备类索引:
live-equipment-index(同步equipment、chassistypes、containertypes) - 用户类索引:
live-users-index(同步drivers、customers、chassisowners、containerowners) - 业务数据索引:
live-loads-index(同步loads)
修改Monstache的[[mapping]]配置,将对应集合映射到新索引,避免单索引的资源竞争。
4. 降低初始同步的并发压力
- 初始同步期间,暂停非必要的搜索查询,或降低查询并发度。
- 进一步降低初始同步的并发数:
direct-read-concur = 1
5. 优化同步脚本性能
- 检查
./src/elasticSyncScripts/下的脚本,避免生成过大的文档(如冗余字段、过深嵌套结构),减少单文档体积。 - 确保脚本无阻塞或耗时操作,防止Monstache积累过多待写入文档。
内容的提问来源于stack exchange,提问作者Vivek Paladiya
相关产品推荐
相关产品推荐

