Redis与Memcached模式下收发数据不一致问题求助(附Redis示例代码)
Hey there, let's dig into this data inconsistency issue you're hitting with Redis and Memcached when using pickle. I'll walk through possible causes and fixes based on the code snippet you shared.
First, fix the obvious code glitch
Looking at your code, the line print(redis_connection.get('foo... is cut off—you probably meant redis_connection.get('foo'), right? If you're using a typo'd key name when retrieving data, that's an instant cause of mismatched results. Assuming that's just a paste error, let's move to more subtle issues.
Possible Cause 1: Pickle protocol version or encoding mismatches
- Pickle uses different protocol versions, and if you serialize with a newer protocol (like v5 introduced in Python 3.8) but try to deserialize in an older Python version, you'll get errors or corrupted data.
- While pickle handles binary data natively, if you're working across systems with different default encodings, there's a tiny chance of hidden string conversion issues—though this is rare since pickle outputs raw bytes.
Possible Cause 2: Data truncation or storage limits
- Your serialized data includes non-ASCII characters (
фывфывфывфа). If your client library treats binary data as strings (instead of raw bytes), it might mangle the data during encoding/decoding. - Memcached has a default 1MB limit per key-value pair; Redis defaults to 512MB. If your pickled data exceeds these limits, it'll get truncated, leading to incomplete or unreadable data when retrieved.
Possible Cause 3: Outdated client libraries or connection issues
- Old versions of
redis-pyor Memcached clients might have bugs handling binary data. For example, some olderredis-pyversions auto-convert bytes to strings by default, which breaks pickled data. - Connection pool issues (like reused connections with leftover data) could also cause partial data writes/reads, though this is less common.
Fixes to try
1. Fix your retrieval code and verify serialization/deserialization
Make sure you're retrieving the correct key and fully deserializing the data:
# Complete the retrieval step retrieved_pickle = redis_connection.get('foo') # Deserialize and compare to original data retrieved_data = pickle.loads(retrieved_pickle) print("Original:", d) print("Retrieved:", retrieved_data)
2. Use a compatible Pickle protocol
Specify a protocol version that works across all your Python environments (protocol 4 is compatible with Python 3.4+):
pickle_dumps = pickle.dumps(d, protocol=4)
3. Ensure clients handle binary data correctly
- For
redis-py, explicitly setencoding=Noneto prevent automatic string conversion:redis_connection = redis.Redis(**REDIS, encoding=None) - For Memcached (using
pymemcache), use the binary protocol to avoid encoding issues:from pymemcache.client.base import Client memcached_client = Client((IGNITE['host'], IGNITE['port'])) memcached_client.set('foo', pickle_dumps) # Pass raw bytes directly
4. Check storage limits
- If your data is larger than Memcached's default 1MB, adjust the
max_item_sizesetting in your Memcached config. - For Redis, confirm the
maxmemorysetting isn't causing your key to be evicted prematurely.
5. Debug with byte-level checks
Compare the length and content of the serialized data before and after storage to pinpoint truncation or corruption:
# Print length of serialized data print("Serialized length:", len(pickle_dumps)) # Retrieve and check length retrieved = redis_connection.get('foo') print("Retrieved length:", len(retrieved)) # Compare raw bytes print("Bytes match:", pickle_dumps == retrieved)
If lengths don't match, data is being truncated. If bytes don't match, the data is being modified during storage/retrieval.
Wrap-up
Start with fixing code typos and verifying client binary handling, then check pickle protocol compatibility and storage limits. These steps should help you track down exactly where the data is getting messed up.
内容的提问来源于stack exchange,提问作者Владимир Чиликов

