如何实现逐个解析序列化Protobuf消息字段的ParseMsgField函数?
问题描述
我熟悉Protobuf的Wire格式,知道它的设计里没在传输中包含序列化消息的大小,也清楚对应的两种处理策略:
- 自行分隔消息:手动通过IO流先读取消息大小再解析
- 使用
google/protobuf/util/delimited_message_util.h工具类处理
但我有个特殊需求:面对无法控制的已序列化消息,希望能逐个加载其中的字段。因为每个字段的ID和大小都已经编码在Wire格式里,理论上可以实现一个类似ParseMsgField的函数。相关的Proto定义如下:
package my_proto; message SubType0 { uint32 Val = 1; } message SubType1 { uint32 Val = 1; } message Root { repeated SubType0 subtype = 1; repeated SubType1 subtype = 2; }
对应的C++示例代码框架如下:
#include "my_proto.pb.h" // 需要实现的函数 int ParseMsgField( google::protobuf::Descriptor* root_descriptor, google::protobuf::Message* field_msg, google::protobuf::io::FileInputStream& fin ); int main() { int fd = open("my_file", O_RDONLY); google::protobuf::io::FileInputStream fin(fd); auto* root_descriptor = my_proto::Root::descriptor(); google::protobuf::Message* field_msg; while(true) { int field_id = ParseMsgField(root_descriptor, field_msg, fin); if (field_id == -1) { break; } // 根据字段ID或类型转换处理每个字段 } }
现在需要解决两个问题:
- 如何实现上述的
ParseMsgField函数? - 这种实现是否要求
Root的所有子字段都是消息类型?
解决方案
一、实现ParseMsgField函数
要实现逐个解析字段的功能,需要直接操作Protobuf的Wire格式编码,结合Protobuf的反射API动态创建和填充消息对象。以下是具体实现:
#include <google/protobuf/descriptor.h> #include <google/protobuf/message.h> #include <google/protobuf/io/coded_stream.h> int ParseMsgField( google::protobuf::Descriptor* root_descriptor, google::protobuf::Message*& field_msg, // 用引用传递,用于输出创建的消息对象 google::protobuf::io::FileInputStream& fin ) { google::protobuf::io::CodedInputStream coded_input(&fin); // 根据实际业务调整字节限制,避免内存溢出 coded_input.SetTotalBytesLimit(INT_MAX, INT_MAX); // 读取字段的Tag(包含字段ID和类型信息) uint32_t tag; if (!coded_input.ReadTag(&tag)) { // 流读取完毕或出错 return -1; } int field_id = google::protobuf::WireFormat::GetTagFieldNumber(tag); auto field_wire_type = google::protobuf::WireFormat::GetTagWireType(tag); // 从Root描述符中查找对应字段 const google::protobuf::FieldDescriptor* field_desc = root_descriptor->FindFieldByNumber(field_id); if (!field_desc) { // 遇到未知字段,跳过对应数据 google::protobuf::WireFormat::SkipField(&coded_input, tag); return field_id; } // 处理消息类型字段,非消息类型字段可扩展逻辑 if (field_desc->type() != google::protobuf::FieldDescriptor::TYPE_MESSAGE) { // 跳过非消息类型字段的数据 google::protobuf::WireFormat::SkipField(&coded_input, tag); return field_id; } // 获取字段对应的消息类型描述符,创建消息实例 const google::protobuf::Descriptor* msg_desc = field_desc->message_type(); const google::protobuf::Message* prototype = google::protobuf::MessageFactory::generated_factory()->GetPrototype(msg_desc); if (!prototype) { // 无法创建消息实例,跳过字段 google::protobuf::WireFormat::SkipField(&coded_input, tag); return field_id; } // 克隆新的消息对象并解析内容 field_msg = prototype->New(); if (!google::protobuf::WireFormat::ParseMessage(&coded_input, field_msg, msg_desc)) { delete field_msg; field_msg = nullptr; return -1; } return field_id; }
实现说明:
- 用
CodedInputStream处理底层Wire格式的IO操作,这是Protobuf提供的原生工具类 - 先读取Tag解析字段ID和类型,再通过Root描述符匹配对应字段
- 利用反射API动态创建消息实例,调用
ParseMessage解析字段内容 - 遇到未知字段或非消息类型字段时,调用
SkipField跳过数据,保证流的位置正确
二、关于子类型是否必须为消息类型的疑问
这种实现不要求Root的所有子类型都是消息类型,但需要根据字段类型扩展处理逻辑:
- 消息类型字段:按上述逻辑创建实例并解析内容
- 基本类型字段(如uint32、string等):可以添加分支逻辑,用
CodedInputStream的ReadVarint32、ReadString等方法读取数据,可通过额外输出参数返回字段值 - 重复字段:Wire格式中重复消息会被编码为多个独立的Tag+数据项,当前函数可以直接逐个读取这些重复项,无需额外修改
如果需要支持非消息类型字段的解析,只需在函数中增加对应类型的处理分支即可。
内容的提问来源于stack exchange,提问作者supernun
相关产品推荐
相关产品推荐

