如何统计MongoDB中各数据库下每个集合的文档数量?
批量统计MongoDB所有数据库和集合的文档数
我来给你一个实用的批量统计方案,用Python的pymongo就能轻松搞定。你不需要手动逐个处理每个集合,只需要通过遍历数据库和集合的方式,就能批量完成文档数统计,下面是具体的实现步骤和代码:
1. 先安装并连接MongoDB客户端
首先确保你已经安装了pymongo(执行pip install pymongo即可),然后建立和本地MongoDB的连接:
from pymongo import MongoClient # 连接本地默认端口的MongoDB,若有认证需求,添加用户名密码即可 client = MongoClient('mongodb://localhost:27017/')
2. 遍历数据库+集合,批量统计文档数
接下来我们遍历所有数据库,再逐个处理每个数据库下的集合(可以选择排除系统集合,比如system.indexes,也可根据需求保留),然后统计每个集合的文档数量:
# 遍历所有数据库 for db_name in client.list_database_names(): db = client[db_name] print(f"\n=== 数据库: {db_name} ===") # 遍历当前数据库下的所有集合 for coll_name in db.list_collection_names(): # 可选:跳过系统集合,若需要统计则注释掉下面的continue if coll_name.startswith('system.'): continue collection = db[coll_name] # 方式1:用count_documents获取精确文档数(适合需要准确值的场景) doc_count = collection.count_documents({}) # 方式2:用你之前尝试的聚合操作统计(效果和上面一致) # doc_count = next(collection.aggregate([{"$group": {"_id": None, "count": {"$sum": 1}}}]))["count"] print(f"集合 {coll_name}: {doc_count} 个文档")
3. 额外优化建议
- 如果你只需要快速估计文档数(不需要绝对精确),可以用
collection.estimated_document_count(),它基于集合元数据计算,速度更快 - 若想把统计结果保存下来,可以存入字典或写入文件,示例如下:
# 用字典存储所有统计结果 stats = {} for db_name in client.list_database_names(): db_stats = {} db = client[db_name] for coll_name in db.list_collection_names(): if not coll_name.startswith('system.'): db_stats[coll_name] = db[coll_name].count_documents({}) stats[db_name] = db_stats # 打印最终统计结果 print(stats)
内容的提问来源于stack exchange,提问作者Rich
相关产品推荐
相关产品推荐

