补丁分析场景下结构体列表对齐方案 匹配游戏新旧版本网络消息结构
游戏网络协议版本自动匹配迁移方案
核心前提
你当前的场景有一个非常关键的稳定特性可以利用:消息和消息类型的整体排列顺序保持不变,这个特性是做自动匹配的核心基础,完全可以基于序列对齐算法实现90%以上的自动迁移,剩下少部分特殊变化人工校验即可。
问题目标
将旧版本已经完成逆向注释的消息结构信息,自动迁移到新版本未注释的原始消息结构中,自动填充大类名称、消息名称、字段名称的注释,新增内容自动标记。
示例参考数据
旧版本已注释消息结构
Authentication: message Login: opcode: 1 fields: - string mail - string password message LoginResponse: opcode: 2 fields: - string token Chat: message ChatSend: opcode: 3 fields: - string channel - string message message ChatReceive: opcode: 4 fields: - string channel - string user - string message
新版本未注释消息结构
Type1: # Authentication message Unk1: # Login opcode: 1 fields: - string unk1 # mail - string unk2 # password _ string unk3 # new field message Unk2: # LoginResponse opcode: 2 fields: - string unk1 # token Type2: # new Type message Unk3: opcode: 3 fields: - Vec3 unk1 - float unk2 Type3: # Chat message Unk4: # ChatSend opcode: 4 fields: - string unk1 # channel - string unk2 # message message Unk5: # new message opcode: 5 fields: - string unk1 - string unk2 message Unk6: # ChatReceive opcode: 6 fields: - string unk1 # channel - string unk2 # user - string unk3 # message
实现思路
基于最长公共子序列(LCS)算法做三层序列对齐:
- 第一层:消息大类序列对齐,匹配旧版本大类列表和新版本大类列表,对齐后的大类直接迁移旧的大类名称,未匹配到的标记为新增大类
- 第二层:单个大类下的消息列表对齐,匹配该大类下旧版本消息列表和新版本消息列表,对齐后的消息结合
opcode差值、字段类型序列相似度做二次校验,校验通过迁移消息名称 - 第三层:单个消息下的字段列表对齐,匹配该消息下旧版本字段列表和新版本字段列表,类型一致的对齐项直接迁移字段名称,未匹配到的标记为新增字段
Python实现代码
from difflib import SequenceMatcher def lcs_match(old_list, new_list, similarity_func): """通用LCS匹配函数,返回匹配对(old_idx, new_idx)和未匹配的新旧索引""" matcher = SequenceMatcher(None, old_list, new_list, autojunk=False) matches = [] for tag, i1, i2, j1, j2 in matcher.get_opcodes(): if tag == 'equal': for i, j in zip(range(i1, i2), range(j1, j2)): if similarity_func(old_list[i], new_list[j]) > 0.7: # 相似度阈值可调整 matches.append((i, j)) old_matched = {i for i, j in matches} new_matched = {j for i, j in matches} return matches, list(set(range(len(old_list))) - old_matched), list(set(range(len(new_list))) - new_matched) def type_similarity(old_type, new_type): """大类相似度判断,优先看消息数量占比,其次看opcode范围重合度""" old_msgs = old_type['messages'] new_msgs = new_type['messages'] old_opcodes = {msg['opcode'] for msg in old_msgs} new_opcodes = {msg['opcode'] for msg in new_msgs} opcode_overlap = len(old_opcodes & new_opcodes) / len(old_opcodes | new_opcodes) if old_opcodes | new_opcodes else 0 count_similarity = min(len(old_msgs), len(new_msgs)) / max(len(old_msgs), len(new_msgs)) if max(len(old_msgs), len(new_msgs)) !=0 else 0 return (opcode_overlap * 0.6 + count_similarity * 0.4) def msg_similarity(old_msg, new_msg): """单条消息相似度判断,优先看opcode差值,其次看字段类型序列相似度""" opcode_diff = abs(old_msg['opcode'] - new_msg['opcode']) opcode_score = 1 if opcode_diff <= 2 else max(0, 1 - opcode_diff * 0.2) old_fields = [f['type'] for f in old_msg['fields']] new_fields = [f['type'] for f in new_msg['fields']] field_matcher = SequenceMatcher(None, old_fields, new_fields, autojunk=False) field_similarity = field_matcher.ratio() return (opcode_score * 0.7 + field_similarity * 0.3) def field_similarity(old_field, new_field): """字段相似度判断,只要类型一致就匹配""" return 1 if old_field['type'] == new_field['type'] else 0 def migrate_annotation(old_protocol, new_protocol): """完整迁移流程""" # 第一步:匹配大类 type_matches, old_unmatched_types, new_unmatched_types = lcs_match( old_protocol['types'], new_protocol['types'], type_similarity ) for old_type_idx, new_type_idx in type_matches: old_type = old_protocol['types'][old_type_idx] new_type = new_protocol['types'][new_type_idx] new_type['name'] = old_type['name'] # 迁移大类名称 # 第二步:匹配该大类下的消息 msg_matches, old_unmatched_msgs, new_unmatched_msgs = lcs_match( old_type['messages'], new_type['messages'], msg_similarity ) for old_msg_idx, new_msg_idx in msg_matches: old_msg = old_type['messages'][old_msg_idx] new_msg = new_type['messages'][new_msg_idx] new_msg['name'] = old_msg['name'] # 迁移消息名称 # 第三步:匹配该消息下的字段 field_matches, old_unmatched_fields, new_unmatched_fields = lcs_match( old_msg['fields'], new_msg['fields'], field_similarity ) for old_field_idx, new_field_idx in field_matches: old_field = old_msg['fields'][old_field_idx] new_field = new_msg['fields'][new_field_idx] new_field['name'] = old_field['name'] # 迁移字段名称 # 标记新增字段 for idx in new_unmatched_fields: new_type['messages'][new_msg_idx]['fields'][idx]['name'] = 'new_field' # 标记新增消息 for idx in new_unmatched_msgs: new_type['messages'][idx]['name'] = 'new_message' # 标记新增大类 for idx in new_unmatched_types: new_protocol['types'][idx]['name'] = 'new_type' return new_protocol # 示例使用 if __name__ == "__main__": # 这里是把旧版本和新版本的消息结构解析成下面的字典格式即可运行 old_protocol = { "types": [ { "name": "Authentication", "messages": [ { "name": "Login", "opcode": 1, "fields": [{"type": "string", "name": "mail"}, {"type": "string", "name": "password"}] }, { "name": "LoginResponse", "opcode": 2, "fields": [{"type": "string", "name": "token"}] } ] }, { "name": "Chat", "messages": [ { "name": "ChatSend", "opcode": 3, "fields": [{"type": "string", "name": "channel"}, {"type": "string", "name": "message"}] }, { "name": "ChatReceive", "opcode": 4, "fields": [{"type": "string", "name": "channel"}, {"type": "string", "name": "user"}, {"type": "string", "name": "message"}] } ] } ] } # 新版本结构同理解析后传入migrate_annotation函数即可
调优建议
- 可以根据你实际的协议变化规律调整相似度的权重和阈值,比如如果opcode基本是连续递增的,可以把opcode的权重进一步提高
- 对于匹配度低于阈值的内容可以输出警告,人工二次校验即可
- 可以加入字段数量、消息长度等额外特征进一步提高匹配准确率
内容的提问来源于stack exchange,提问作者ACB
相关产品推荐
相关产品推荐

