Java写入多Protobuf消息至文件,Python端如何反序列化读取
writeDelimitedTo/parseDelimitedFrom) Great question! I’ve run into this exact issue before—Python’s Protobuf library doesn’t include built-in delimited read/write methods like Java does, but we can easily replicate that behavior by manually adding length prefixes (using varint encoding, which is what Java uses under the hood). Here’s how to do it:
Writing Multiple Messages to a File
To write multiple messages safely, each message needs to be prefixed with its serialized length (encoded as a varint). This tells the reader exactly how many bytes to read for each individual message, preventing parsing errors when dealing with multiple entries.
Here’s a reusable function to handle this:
import google.protobuf.internal.encoder as encoder import google.protobuf.internal.decoder as decoder def write_delimited_to(message, file): # Serialize the message to raw bytes serialized_msg = message.SerializeToString() # Write the varint-encoded length of the serialized message encoder._VarintBytes(len(serialized_msg), file) # Write the actual serialized message data file.write(serialized_msg) # Example usage: # from your_proto_def import YourMessageType # with open("messages.bin", "wb") as f: # msg1 = YourMessageType(data="first message") # write_delimited_to(msg1, f) # msg2 = YourMessageType(data="second message") # write_delimited_to(msg2, f)
Reading Multiple Messages from a File
To read back the messages, we first read the varint length prefix, then read exactly that number of bytes, and finally parse those bytes into a Protobuf message. We can loop this process until we reach the end of the file.
Here’s the corresponding read function:
def parse_delimited_from(message_type, file): while True: try: # Read the varint length prefix (max 10 bytes for varint) length_bytes = file.read(10) if not length_bytes: break # End of file reached length, _ = decoder._DecodeVarint(length_bytes, 0) # Read exactly the number of bytes specified by the length serialized_msg = file.read(length) if not serialized_msg: break # Parse and yield the message message = message_type() message.ParseFromString(serialized_msg) yield message except EOFError: break # Example usage: # with open("messages.bin", "rb") as f: # for msg in parse_delimited_from(YourMessageType, f): # print(msg.data)
Key Notes
- Cross-Language Compatibility: This method matches Java’s
writeDelimitedTobehavior perfectly, so files written with this Python code can be read by Java (and vice versa) without any issues. - Error Handling: The example includes basic EOF handling, but you might want to add extra checks (like handling invalid length values or corrupted serialized data) depending on your use case.
- Performance: Varint encoding is compact and efficient, and reading exactly the required bytes avoids unnecessary buffering, making this approach fast even for large datasets.
内容的提问来源于stack exchange,提问作者user3162649

