无需Protobuf解码Ignition发送的MQTT Sparkplug B消息?
SparkplugB v1.0 MQTT消息解码难题
我刚接触MQTT,完全不懂Google Protobuf。接手了一个无文档的Python项目,现在卡在解码这一步。跟着Paho MQTT官方教程写了脚本,能收到消息但解不出来。试了各种编码,只有latin-1能出点可读内容,用utf-8解码直接报错:can't decode byte 0xf3: invalid continuation type。
消息来自Ignition服务器,用的是SparkplugB v1.0格式。之前的项目用Protobuf解码,但我不想直接加那600多行代码,想先评估必要性。后来发现新的payload里有十六进制头和明文数据,感觉可能必须用Protobuf才能完整解码。
现有代码
import paho.mqtt.client as mqtt import json def decodeMQTT(msg): # # it's coming in as SparkplugB v1.0 # print('x!x!x-- decoding payload... --x!x!x') def on_connect(client, userdata, flags, rc): print('connected with result code: ' + str(rc)) client.subscribe("spBv1.0/spanky/+/TCG4 Test1/#") def on_message(client, userdata, msg): tokens = msg.topic.split('/') print('--- tokens ---> ',tokens) # OUTPUTS: --- tokens ---> ['spBv1.0', 'spanky', 'NCMD', 'TCG4 Test1'] print('-+-+-+- raw payload: ', msg.payload) # OUTPUTS: -+-+-+- raw payload: b'\x08\xa8\xbe\xbc\xad\x9f1\x12\x12\x10\x03\x18\xa8\xbe\xbc\xad\x9f1 # \t8\x00e\x00\x00\xa0A\x18\xff\xff\xff\xff\xff\xff\xff\xff\xff\x01' try: # # this should be using UTF-8 to decode the message string... but it isn't # which leads me to believe it has something to do with protobuf (dammit) # if UTF-8 is used, i get this error: # »» can't decode byte 0xf3: invalid continuation type «« # which i can't seem to find any explanation for... :P # decoded = msg.payload.decode('latin-1') print('♥♥♥-» decoded payload: ', str(decoded), '«-♥♥♥') # # this prints: ♥♥♥-» decoded payload: ÅÒÖ1↕& # ‼/tags/inputs/input1►♥↑ÇÒÖ1 ♀z♦2.01↑^ «-♥♥♥ # NOTE the line break is part of the decoded output # except Exception as ex: print('↓↓↓↓↓ decoding failed: ', str(ex), '↓↓↓↓↓') def main(): client = mqtt.Client() client.on_connect = on_connect client.on_message = on_message client.username_pw_set('fnork', 'blarp') client.connect("nunyabiznitz", 16500, 10) client.loop_forever() if __name__ == '__main__': main()
后续收到的Payload示例
\x08\xd0\xca\x99\xf4\x9f1\x12&\n\x13/tags/inputs/input1\x10\x03\x18\x8e\xcd\x99\xf4\x9f1 \x0cz\x042.01\x18\xae\x01
解决方案
结论:必须用Protobuf解码
SparkplugB v1.0是基于Protobuf的二进制协议,不是纯文本格式。你用latin-1看到的零星可读内容,只是Payload里碰巧和字符编码重合的字节,核心的结构化数据全被Protobuf二进制包着,不用Protobuf根本解不出来。
不用手动加600行代码,用官方工具生成解析类
旧项目里的600多行代码应该是手写的Protobuf解析逻辑,完全没必要用——用官方工具生成的代码更规范,步骤如下:
- 拿到SparkplugB v1.0的标准Protobuf定义文件
sparkplug_b.proto(内容是固定的公开标准) - 用Protobuf编译器
protoc生成Python解析类:protoc --python_out=. sparkplug_b.proto - 在你的代码里导入生成的类,直接解码Payload
修改后的on_message函数示例:
import sparkplug_b_pb2 # 导入生成的Protobuf类 def on_message(client, userdata, msg): tokens = msg.topic.split('/') print('--- tokens ---> ',tokens) print('-+-+-+- raw payload: ', msg.payload) try: # 解码SparkplugB消息 sp_payload = sparkplug_b_pb2.Payload() sp_payload.ParseFromString(msg.payload) # 打印解析后的结构化数据 print('=== 解析完成 ===') print(f'时间戳: {sp_payload.timestamp}') for metric in sp_payload.metrics: # 根据数据类型取对应的值,这里以float和string为例 if metric.datatype == 3: metric_value = metric.float_value elif metric.datatype == 12: metric_value = metric.string_value else: metric_value = f"未处理类型: {metric.datatype}" print(f'指标名: {metric.name}, 值: {metric_value}') except Exception as ex: print('↓↓↓↓↓ 解码失败: ', str(ex), '↓↓↓↓↓')
为什么utf-8会报错?
utf-8对字节序列有严格规则,比如0xf3是多字节字符的起始字节,后面必须跟特定范围的续字节,但Protobuf的二进制数据是结构化的字节流,完全不符合utf-8的编码规则,所以解码报错是必然的。
内容的提问来源于stack exchange,提问作者WhiteRau
相关产品推荐
相关产品推荐

