You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MongoDB分片数据分布不均问题排查与解决请求

MongoDB分片数据不均问题解决方案

问题背景

我有一个存储80万条文档的MongoDB数据库,部署了13台shard服务器以提升访问速度。为每个文档分配了一个字母作为shards字段值,各字母对应的文档数量均匀(约5万条/字母),并将该字段设为哈希分片键,执行sh.shardCollection("test.testCollection", { "shards": "hashed" })进行分片。但数据仅分布在两台shard服务器上(a节点占28%,k节点占72%),其余11台无数据,需要将数据均匀分配到全部13台shard节点。

环境信息

Shard节点详情

[
  {
    _id: 'a',
    host: 'a/127.0.0.1:21000,127.0.0.1:21001,127.0.0.1:21002',
    state: 1,
    topologyTime: Timestamp({ t: 1675107083, i: 3 })
  },
  {
    _id: 'b',
    host: 'b/127.0.0.1:22000,127.0.0.1:22001,127.0.0.1:22002',
    state: 1,
    topologyTime: Timestamp({ t: 1675107100, i: 5 })
  },
  {
    _id: 'c',
    host: 'c/127.0.0.1:23000,127.0.0.1:23001,127.0.0.1:23002',
    state: 1,
    draining: true
  },
  {
    _id: 'd',
    host: 'd/127.0.0.1:23010,127.0.0.1:23011,127.0.0.1:23012',
    state: 1,
    topologyTime: Timestamp({ t: 1676821653, i: 5 })
  },
  {
    _id: 'e',
    host: 'e/127.0.0.1:23020,127.0.0.1:23021,127.0.0.1:23022',
    state: 1,
    topologyTime: Timestamp({ t: 1676821663, i: 5 })
  },
  {
    _id: 'f',
    host: 'f/127.0.0.1:23030,127.0.0.1:23031,127.0.0.1:23032',
    state: 1,
    topologyTime: Timestamp({ t: 1676821668, i: 1 })
  },
  {
    _id: 'g',
    host: 'g/127.0.0.1:23040,127.0.0.1:23041,127.0.0.1:23042',
    state: 1,
    topologyTime: Timestamp({ t: 1676821673, i: 5 })
  },
  {
    _id: 'h',
    host: 'h/127.0.0.1:23050,127.0.0.1:23051,127.0.0.1:23052',
    state: 1,
    topologyTime: Timestamp({ t: 1676821678, i: 5 })
  },
  {
    _id: 'j',
    host: 'j/127.0.0.1:23060,127.0.0.1:23061,127.0.0.1:23062',
    state: 1,
    topologyTime: Timestamp({ t: 1676821685, i: 5 })
  },
  {
    _id: 'k',
    host: 'k/127.0.0.1:23070,127.0.0.1:23071,127.0.0.1:23072',
    state: 1,
    topologyTime: Timestamp({ t: 1676821689, i: 5 })
  },
  {
    _id: 'l',
    host: 'l/127.0.0.1:23080,127.0.0.1:23081,127.0.0.1:23082',
    state: 1,
    topologyTime: Timestamp({ t: 1676821694, i: 5 })
  },
  {
    _id: 'm',
    host: 'm/127.0.0.1:23090,127.0.0.1:23091,127.0.0.1:23092',
    state: 1,
    topologyTime: Timestamp({ t: 1676821698, i: 5 })
  },
  {
    _id: 'n',
    host: 'n/127.0.0.1:24000,127.0.0.1:24001,127.0.0.1:24002',
    state: 1,
    topologyTime: Timestamp({ t: 1676821708, i: 4 })
  }
]

当前数据分布

Shard a at a/127.0.0.1:21000,127.0.0.1:21001,127.0.0.1:21002
{
  data: '125.57MiB',
  docs: 227420,
  chunks: 1,
  'estimated data per chunk': '125.57MiB',
  'estimated docs per chunk': 227420
}
Shard k at k/127.0.0.1:23070,127.0.0.1:23071,127.0.0.1:23072
{
  data: '326.31MiB',
  docs: 576209,
  chunks: 1,
  'estimated data per chunk': '326.31MiB',
  'estimated docs per chunk': 576209
}

文档示例

{
  "_id": {
    "$oid": "63dd7324289226c918818c55"
  },
  "Title": "",
  "Product": {
    "web1": {
      "Harry Potter and the Chamber of Secrets: 2/7 (Harry Potter 2)": {
        "Price": 15,
        "Url": "https://www.amazon.com/Harry-Potter-Chamber-Secrets-Book/dp/B017V4IPPO/ref=sr_1_2?crid=GCT8C7Z3Q4SE&keywords=Harry+Potter+and+the+Chamber+of+Secrets&qid=1676836656&sprefix=harry+potter+and+the+chamber+of+secrets%2Caps%2C230&sr=8-2",
        "Time": {
          "$date": {
            "$numberLong": "1676669514749"
          }
        }
      }
    }
  },
  "Category": [
    "Book",
    "Fantasy"
  ],
  "Time": {
    "$date": {
      "$numberLong": "1676669514749"
    }
  },
  "shards": "h"
}

解决方案

1. 修复异常shard节点

首先处理标记为draining: true的c节点,若该节点需要参与分片,执行以下命令恢复可用状态:

sh.stopBalancer()
use admin
// 完成节点退役流程
db.runCommand( { removeShard: "c" } )
// 重新添加节点
sh.addShard("c/127.0.0.1:23000,127.0.0.1:23001,127.0.0.1:23002")
sh.startBalancer()

执行sh.status()确认所有shard节点状态为state:1(可用)。

2. 手动拆分大chunk

当前集合仅2个chunk,无法自动分配到13个节点,需拆分出足够数量的chunk:

// 查看当前chunk范围
use config
db.chunks.find({ns: "test.testCollection"}).sort({min:1})

// 自动拆分出26个chunk(冗余设计,确保每个shard至少分配2个)
sh.splitCollection("test.testCollection", { shards: "hashed" }, { numChunks: 26 })

若自动拆分效果不佳,可针对每个字母对应的哈希范围手动拆分:

sh.splitFind("test.testCollection", { shards: "a" })
sh.splitFind("test.testCollection", { shards: "b" })
// 依次对所有字母执行拆分

3. 触发chunk均衡

确保平衡器运行,MongoDB会自动将chunk迁移到空闲shard:

sh.startBalancer()

若平衡器未自动触发,手动迁移chunk到空闲节点:

use config
// 获取所有空闲shard的ID
var emptyShards = db.shards.find({_id: {$nin: ["a","k"]}}).map(s => s._id)

// 从shard k迁移chunk到空闲节点
var chunksFromK = db.chunks.find({ns: "test.testCollection", shard: "k"}).limit(emptyShards.length)
chunksFromK.forEach(function(chunk) {
  var target = emptyShards.shift()
  sh.moveChunk("test.testCollection", chunk.min, target)
})

// 从shard a迁移剩余chunk
var remainingShards = db.shards.find({_id: {$nin: ["a","k", ...emptyShards]}}).map(s => s._id)
var chunksFromA = db.chunks.find({ns: "test.testCollection", shard: "a"}).limit(remainingShards.length)
chunksFromA.forEach(function(chunk) {
  var target = remainingShards.shift()
  sh.moveChunk("test.testCollection", chunk.min, target)
})

4. 验证均衡效果

执行以下命令确认数据分布:

use test
db.testCollection.getShardDistribution()
sh.status()

目标是每个shard节点的文档数量接近80万/13≈6.15万条。

5. 长期优化措施

  • 分片后再批量插入数据:避免分片前批量插入导致chunk集中在少数节点
  • 调整chunk大小:根据数据量调整默认64MB的chunk大小,确保自动拆分触发
use config
db.settings.updateOne({_id:"chunksize"}, {$set:{value: 32}}, {upsert:true})
  • 监控分片状态:定期执行sh.status()和getShardDistribution(),确保数据分布均匀

内容的提问来源于stack exchange,提问作者BayGold

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:00:42