MongoDB分片数据分布不均问题排查与解决请求
MongoDB分片数据不均问题解决方案
问题背景
我有一个存储80万条文档的MongoDB数据库,部署了13台shard服务器以提升访问速度。为每个文档分配了一个字母作为shards字段值,各字母对应的文档数量均匀(约5万条/字母),并将该字段设为哈希分片键,执行sh.shardCollection("test.testCollection", { "shards": "hashed" })进行分片。但数据仅分布在两台shard服务器上(a节点占28%,k节点占72%),其余11台无数据,需要将数据均匀分配到全部13台shard节点。
环境信息
Shard节点详情
[ { _id: 'a', host: 'a/127.0.0.1:21000,127.0.0.1:21001,127.0.0.1:21002', state: 1, topologyTime: Timestamp({ t: 1675107083, i: 3 }) }, { _id: 'b', host: 'b/127.0.0.1:22000,127.0.0.1:22001,127.0.0.1:22002', state: 1, topologyTime: Timestamp({ t: 1675107100, i: 5 }) }, { _id: 'c', host: 'c/127.0.0.1:23000,127.0.0.1:23001,127.0.0.1:23002', state: 1, draining: true }, { _id: 'd', host: 'd/127.0.0.1:23010,127.0.0.1:23011,127.0.0.1:23012', state: 1, topologyTime: Timestamp({ t: 1676821653, i: 5 }) }, { _id: 'e', host: 'e/127.0.0.1:23020,127.0.0.1:23021,127.0.0.1:23022', state: 1, topologyTime: Timestamp({ t: 1676821663, i: 5 }) }, { _id: 'f', host: 'f/127.0.0.1:23030,127.0.0.1:23031,127.0.0.1:23032', state: 1, topologyTime: Timestamp({ t: 1676821668, i: 1 }) }, { _id: 'g', host: 'g/127.0.0.1:23040,127.0.0.1:23041,127.0.0.1:23042', state: 1, topologyTime: Timestamp({ t: 1676821673, i: 5 }) }, { _id: 'h', host: 'h/127.0.0.1:23050,127.0.0.1:23051,127.0.0.1:23052', state: 1, topologyTime: Timestamp({ t: 1676821678, i: 5 }) }, { _id: 'j', host: 'j/127.0.0.1:23060,127.0.0.1:23061,127.0.0.1:23062', state: 1, topologyTime: Timestamp({ t: 1676821685, i: 5 }) }, { _id: 'k', host: 'k/127.0.0.1:23070,127.0.0.1:23071,127.0.0.1:23072', state: 1, topologyTime: Timestamp({ t: 1676821689, i: 5 }) }, { _id: 'l', host: 'l/127.0.0.1:23080,127.0.0.1:23081,127.0.0.1:23082', state: 1, topologyTime: Timestamp({ t: 1676821694, i: 5 }) }, { _id: 'm', host: 'm/127.0.0.1:23090,127.0.0.1:23091,127.0.0.1:23092', state: 1, topologyTime: Timestamp({ t: 1676821698, i: 5 }) }, { _id: 'n', host: 'n/127.0.0.1:24000,127.0.0.1:24001,127.0.0.1:24002', state: 1, topologyTime: Timestamp({ t: 1676821708, i: 4 }) } ]
当前数据分布
Shard a at a/127.0.0.1:21000,127.0.0.1:21001,127.0.0.1:21002 { data: '125.57MiB', docs: 227420, chunks: 1, 'estimated data per chunk': '125.57MiB', 'estimated docs per chunk': 227420 }
Shard k at k/127.0.0.1:23070,127.0.0.1:23071,127.0.0.1:23072 { data: '326.31MiB', docs: 576209, chunks: 1, 'estimated data per chunk': '326.31MiB', 'estimated docs per chunk': 576209 }
文档示例
{ "_id": { "$oid": "63dd7324289226c918818c55" }, "Title": "", "Product": { "web1": { "Harry Potter and the Chamber of Secrets: 2/7 (Harry Potter 2)": { "Price": 15, "Url": "https://www.amazon.com/Harry-Potter-Chamber-Secrets-Book/dp/B017V4IPPO/ref=sr_1_2?crid=GCT8C7Z3Q4SE&keywords=Harry+Potter+and+the+Chamber+of+Secrets&qid=1676836656&sprefix=harry+potter+and+the+chamber+of+secrets%2Caps%2C230&sr=8-2", "Time": { "$date": { "$numberLong": "1676669514749" } } } } }, "Category": [ "Book", "Fantasy" ], "Time": { "$date": { "$numberLong": "1676669514749" } }, "shards": "h" }
解决方案
1. 修复异常shard节点
首先处理标记为draining: true的c节点,若该节点需要参与分片,执行以下命令恢复可用状态:
sh.stopBalancer() use admin // 完成节点退役流程 db.runCommand( { removeShard: "c" } ) // 重新添加节点 sh.addShard("c/127.0.0.1:23000,127.0.0.1:23001,127.0.0.1:23002") sh.startBalancer()
执行sh.status()确认所有shard节点状态为state:1(可用)。
2. 手动拆分大chunk
当前集合仅2个chunk,无法自动分配到13个节点,需拆分出足够数量的chunk:
// 查看当前chunk范围 use config db.chunks.find({ns: "test.testCollection"}).sort({min:1}) // 自动拆分出26个chunk(冗余设计,确保每个shard至少分配2个) sh.splitCollection("test.testCollection", { shards: "hashed" }, { numChunks: 26 })
若自动拆分效果不佳,可针对每个字母对应的哈希范围手动拆分:
sh.splitFind("test.testCollection", { shards: "a" }) sh.splitFind("test.testCollection", { shards: "b" }) // 依次对所有字母执行拆分
3. 触发chunk均衡
确保平衡器运行,MongoDB会自动将chunk迁移到空闲shard:
sh.startBalancer()
若平衡器未自动触发,手动迁移chunk到空闲节点:
use config // 获取所有空闲shard的ID var emptyShards = db.shards.find({_id: {$nin: ["a","k"]}}).map(s => s._id) // 从shard k迁移chunk到空闲节点 var chunksFromK = db.chunks.find({ns: "test.testCollection", shard: "k"}).limit(emptyShards.length) chunksFromK.forEach(function(chunk) { var target = emptyShards.shift() sh.moveChunk("test.testCollection", chunk.min, target) }) // 从shard a迁移剩余chunk var remainingShards = db.shards.find({_id: {$nin: ["a","k", ...emptyShards]}}).map(s => s._id) var chunksFromA = db.chunks.find({ns: "test.testCollection", shard: "a"}).limit(remainingShards.length) chunksFromA.forEach(function(chunk) { var target = remainingShards.shift() sh.moveChunk("test.testCollection", chunk.min, target) })
4. 验证均衡效果
执行以下命令确认数据分布:
use test db.testCollection.getShardDistribution() sh.status()
目标是每个shard节点的文档数量接近80万/13≈6.15万条。
5. 长期优化措施
- 分片后再批量插入数据:避免分片前批量插入导致chunk集中在少数节点
- 调整chunk大小:根据数据量调整默认64MB的chunk大小,确保自动拆分触发
use config db.settings.updateOne({_id:"chunksize"}, {$set:{value: 32}}, {upsert:true})
- 监控分片状态:定期执行
sh.status()和getShardDistribution(),确保数据分布均匀
内容的提问来源于stack exchange,提问作者BayGold
相关产品推荐
相关产品推荐

