PyMongo游标迭代启动耗时过长问题技术求助
Hey there! That 7-second delay when starting to iterate your MongoDB cursor is brutal, especially since this function gets called often for short-term analysis. Let’s break down the most likely fixes to speed this up:
1. Make Sure You’re Using Proper Indexes
This is the #1 culprit for slow cursor iteration. MongoDB returns cursors quickly, but the actual data fetch happens on the first iteration—and if your query isn’t using an index, it’s doing a full collection scan under the hood.
- First, identify the filter criteria in your
find()call (I see you’re querying for a specificcoin). If you’re sorting by a timestamp field (super common for candlestick data), create a compound index that matches your query and sort order:# Run this once in your MongoDB shell or via Python db.candlestick1m.createIndex({"coin": 1, "timestamp": -1}) - Verify the index is being used with
explain():
Look forlastKlines = db.candlestick1m.find({"coin": coin}).sort("timestamp", -1) print(lastKlines.explain("executionStats"))executionStats.totalDocsExamined—if this number is way higher thanexecutionStats.nReturned, your query isn’t using an index efficiently.
2. Fetch Only the Data You Need
If your find() call is returning full candlestick documents but you only need a few fields (like the latest timestamp), use a projection to reduce data transfer and processing:
# Example: Only fetch the timestamp field, exclude _id lastKlines = db.candlestick1m.find( {"coin": coin}, {"timestamp": 1, "_id": 0} ).sort("timestamp", -1)
This cuts down on the amount of data MongoDB sends over the wire, which can drastically speed up iteration.
3. Adjust Cursor Batch Size
MongoDB fetches documents in batches (default is 101 documents or 16MB). If you’re dealing with a large result set, increasing the batch size can reduce network round trips:
lastKlines = db.candlestick1m.find({"coin": coin}).batch_size(1000)
Note: If your result set is small (like just the latest candlestick), this might not make a big difference, but it’s worth testing.
4. Check Read Preferences (If Using a Replica Set)
If you’re querying a MongoDB replica set, your read preference might be pointing to a secondary node with stale data or higher latency. Try forcing reads from the primary node temporarily to see if that helps:
from pymongo import ReadPreference lastKlines = db.candlestick1m.find({"coin": coin}).read_preference(ReadPreference.PRIMARY)
Just keep in mind this adds load to your primary, so balance that against your performance needs.
5. Cache Repeated Queries (If Possible)
Since this function is called frequently, if you’re querying the same coin multiple times in a short window, cache the latest result in memory (with a short expiration, since crypto data is real-time):
from functools import lru_cache import time # Cache results for 5 seconds (adjust based on your data update frequency) @lru_cache(maxsize=32) def get_last_klines_timestamp(coin): last_kline = db.candlestick1m.find({"coin": coin}, {"timestamp":1, "_id":0}).sort("timestamp", -1).limit(1).next() return last_kline["timestamp"] # In your main function current_timestamp = get_last_klines_timestamp(coin)
This avoids hitting the database entirely for repeated calls within the cache window.
内容的提问来源于stack exchange,提问作者Finrod

