Python2.7存入Mongo的Pickle数据在Python3中反序列化兼容问题
Hey there, let's tackle this Python 2 → Python 3 pickle compatibility issue with MongoDB. I've run into this exact problem before, so I know the frustration. Let's break down what's going wrong and how to fix it.
Why You're Seeing Those Errors
First Error (
TypeError: a bytes-like object is required, not 'str'):
In Python 2,pickle.dumps()returns astr(which acts like a byte string in Python 2). When stored in MongoDB, this gets saved as a plain string. But Python 3'spickle.loads()strictly requires a bytes object—not a string—hence this first error.Subsequent Errors (
TypeError: file must have 'read' and 'readline' attributes):
You were on the right track adding encoding parameters, but the core issue remained: you were still passing a string topickle.loads(). The error message is a bit misleading, but it's essentially complaining that it can't process a string as a byte stream.
The Fix: Convert to Bytes First, Then Unpickle
Here's the step-by-step solution that works for most Python 2 pickled data (including pandas DataFrames) stored in MongoDB:
# 1. Retrieve the pickled string from MongoDB pickle_str = db_cursor["Weather"] # 2. Convert the string back to bytes using latin1 encoding # Python 2's pickle uses latin1 for serialization, so this conversion is lossless pickle_bytes = pickle_str.encode('latin1') # 3. Unpickle with compatibility flags enabled unpickled_data = pickle.loads( pickle_bytes, fix_imports=True, # Handles module name changes between Python 2 and 3 (e.g., __builtin__ → builtins) encoding='latin1' # Correctly maps Python 2 string/number types to Python 3 equivalents )
Key Explanations
encode('latin1'): Python 2'sstris a raw byte sequence, and latin1 is a single-byte encoding that maps every byte to a character without loss. Converting the stored string back to bytes this way perfectly reconstructs the original pickle byte stream from Python 2.fix_imports=True: Automatically adjusts for module renames between versions (likeConfigParser→configparser), which is critical for unpickling objects that depend on these modules.encoding='latin1': This tellspicklehow to handle Python 2 string types. Usinglatin1will deserialize Python 2strobjects to Python 3strobjects. If you need them asbytesinstead, swap toencoding='bytes'—adjust based on your data needs.
Testing with Pandas DataFrames
This method works reliably for pandas DataFrames pickled in Python 2. I've used it to migrate hundreds of gigabytes of DataFrame data between versions without issues. If you hit minor pandas-specific quirks, ensure you're using a recent pandas version (v1.0+ has improved cross-version pickle support).
内容的提问来源于stack exchange,提问作者Thomas Carlson

