Pymongo Cursor迭代瓶颈求助:Geo查询返回Cursor转list耗时过高
list(nearest)耗时高本质不是Python类型转换的开销,核心是MongoDB地理查询执行、结果批量拉取的开销,可按以下优先级优化:
- 优先添加2dsphere地理索引
无索引时地理查询会触发全表扫描,性能极差,给location字段创建2dsphere索引即可将查询耗时降低1~2个数量级,建索引仅需执行一次:
优化后可执行collection = self.database_objs[common_models.ObjModel().current_geographic_collection] collection.create_index([("location", "2dsphere")])print(nearest.explain()["executionStats"]["executionStages"]["stage"]),输出为IXSCAN即代表索引生效。 - 限制返回字段,减少无效数据传输
给find方法添加projection参数,仅返回业务需要的字段,避免拉取冗余内容:nearest = collection.find( {"location": {"$geoWithin": {"$centerSphere": [start, self.distance_radians(self.feet_meter(radius))]}}}, # 按需填写需要返回的字段,不需要的字段设为0即可 projection={"_id": 1, "location": 1, "business_field": 1} ) - 限制返回文档数量
若业务不需要全量匹配结果,添加limit限制返回条数,大幅降低数据拉取耗时:# 比如仅返回最近的100条匹配结果 nearest = collection.find( {"location": {"$geoWithin": {"$centerSphere": [start, self.distance_radians(self.feet_meter(radius))]}}} ).limit(100) - 开启索引覆盖查询
若业务需要的所有字段都包含在索引中,可以创建复合2dsphere索引,实现覆盖查询,无需回表读取原始文档,性能可再提升3~5倍:# 假如你需要返回location、business_field、_id三个字段,创建对应复合索引 collection.create_index([("location", "2dsphere"), ("business_field", 1), ("_id", 1)]) - 调整游标批量拉取参数
匹配结果较多时,调整batch_size参数减少客户端和MongoDB的网络往返次数:nearest = collection.find( {"location": {"$geoWithin": {"$centerSphere": [start, self.distance_radians(self.feet_meter(radius))]}}} ).batch_size(1000) - 超大数据集分片优化
若单集合数据量超过千万级,可按照地理位置字段做分片,将查询范围缩小到单个或少数分片,大幅降低查询扫描的数据量。
内容的提问来源于stack exchange,提问作者Martin Mashalov
相关产品推荐
相关产品推荐

