如何用Python校验XML票据与XML布局的标签顺序及兼容性?
问题
作为开发新手,我正在编写Python脚本校验XML票据(Fiscal Note)与XML布局(Layout)的兼容性。目前已用xml.etree.ElementTree实现了标签相似性校验,现在需要升级脚本,实现标签顺序校验,同时兼容以下场景:
- 票据缺少布局中的某些标签
- 票据多出布局中没有的标签
现有代码如下:
import xml.etree.ElementTree as ET # Function to calculate similarity between two sets def calculate_similarity(set1, set2): intersection = set1.intersection(set2) smallest_set_size = min(len(set1), len(set2)) if smallest_set_size == 0: return 0 similarity = len(intersection) / smallest_set_size return similarity # Function to validate a fiscal note against the layouts def validate_fiscal_note(fiscal_note_xml, similarity_threshold=0.1): fiscal_note_tree = ET.parse(fiscal_note_xml) fiscal_note_root = fiscal_note_tree.getroot() fiscal_note_tags = {child.tag for child in fiscal_note_root.iter()} best_match_code = None best_match_similarity = 0 best_match_code_list = [] for code, layout_tags in layouts_dict.items(): similarity = calculate_similarity(layout_tags, fiscal_note_tags) best_match_code_list.append({'code': code, 'similarity': similarity}) sorted_list = sorted(best_match_code_list, key=lambda x: x['similarity']) if not sorted_list: return 'Any layout was finded' else: return sorted_list # Parse the layouts XML layouts_tree = ET.parse('layouts.xml') layouts_root = layouts_tree.getroot() # Build a dictionary of layouts layouts_dict = {} for layout in layouts_root: layout_code = layout.attrib['code'] # Make sure this matches your XML structure; might need to adjust tags = {child.tag for child in layout} layouts_dict[layout_code] = tags # Validate the fiscal note and print the result result = validate_fiscal_note('Nota 2.xml') print(result[-5:])
解决方案
1. 重构布局字典:存储标签顺序而非集合
原代码用集合存储标签丢失了顺序信息,需要改为存储标签顺序列表:
# 解析布局XML并构建字典(存储标签顺序) layouts_tree = ET.parse('layouts.xml') layouts_root = layouts_tree.getroot() layouts_dict = {} for layout in layouts_root: layout_code = layout.attrib['code'] # 按XML中的实际顺序收集所有标签(含子节点),若只需直接子节点则替换为 `layout` tag_sequence = [child.tag for child in layout.iter()] layouts_dict[layout_code] = tag_sequence
2. 实现顺序匹配算法:最长公共子序列(LCS)
LCS算法可以兼容票据缺标签、多标签的场景,只关注两者共有标签的顺序匹配度:
def calculate_order_match_score(layout_sequence, note_sequence): m, n = len(layout_sequence), len(note_sequence) # 构建动态规划表 dp = [[0]*(n+1) for _ in range(m+1)] for i in range(1, m+1): for j in range(1, n+1): if layout_sequence[i-1] == note_sequence[j-1]: dp[i][j] = dp[i-1][j-1] + 1 else: dp[i][j] = max(dp[i-1][j], dp[i][j-1]) # 顺序得分 = LCS长度 / 布局标签总数(保证布局顺序的优先级) return dp[m][n] / m if m != 0 else 0.0
3. 整合相似性与顺序得分,计算综合匹配度
可根据业务需求调整两者的权重(比如各占50%):
def calculate_composite_score(layout_tags, note_tags, layout_seq, note_seq, sim_weight=0.5, order_weight=0.5): sim_score = calculate_similarity(layout_tags, note_tags) order_score = calculate_order_match_score(layout_seq, note_seq) return sim_weight * sim_score + order_weight * order_score
4. 重构校验函数,同时输出多维度得分
更新校验逻辑,同时计算相似性、顺序匹配度和综合得分:
def validate_fiscal_note(fiscal_note_xml, sim_weight=0.5, order_weight=0.5): fiscal_note_tree = ET.parse(fiscal_note_xml) fiscal_note_root = fiscal_note_tree.getroot() # 收集票据的标签集合和顺序序列(若只需直接子节点则替换为 `fiscal_note_root`) note_tags = {child.tag for child in fiscal_note_root.iter()} note_sequence = [child.tag for child in fiscal_note_root.iter()] match_results = [] for code, layout_seq in layouts_dict.items(): layout_tags = set(layout_seq) sim_score = calculate_similarity(layout_tags, note_tags) order_score = calculate_order_match_score(layout_seq, note_sequence) composite_score = calculate_composite_score(layout_tags, note_tags, layout_seq, note_sequence, sim_weight, order_weight) match_results.append({ 'code': code, 'similarity_score': round(sim_score, 3), 'order_match_score': round(order_score, 3), 'composite_score': round(composite_score, 3) }) # 按综合得分降序排序 sorted_results = sorted(match_results, key=lambda x: x['composite_score'], reverse=True) return sorted_results if sorted_results else 'No matching layout found'
5. 完整整合代码
import xml.etree.ElementTree as ET def calculate_similarity(set1, set2): intersection = set1.intersection(set2) smallest_set_size = min(len(set1), len(set2)) return len(intersection) / smallest_set_size if smallest_set_size != 0 else 0.0 def calculate_order_match_score(layout_sequence, note_sequence): m, n = len(layout_sequence), len(note_sequence) dp = [[0]*(n+1) for _ in range(m+1)] for i in range(1, m+1): for j in range(1, n+1): if layout_sequence[i-1] == note_sequence[j-1]: dp[i][j] = dp[i-1][j-1] + 1 else: dp[i][j] = max(dp[i-1][j], dp[i][j-1]) return dp[m][n] / m if m != 0 else 0.0 def calculate_composite_score(layout_tags, note_tags, layout_seq, note_seq, sim_weight=0.5, order_weight=0.5): sim_score = calculate_similarity(layout_tags, note_tags) order_score = calculate_order_match_score(layout_seq, note_seq) return sim_weight * sim_score + order_weight * order_score def validate_fiscal_note(fiscal_note_xml, sim_weight=0.5, order_weight=0.5): fiscal_note_tree = ET.parse(fiscal_note_xml) fiscal_note_root = fiscal_note_tree.getroot() # 若只需校验根节点直接子节点,替换为: # note_tags = {child.tag for child in fiscal_note_root} # note_sequence = [child.tag for child in fiscal_note_root] note_tags = {child.tag for child in fiscal_note_root.iter()} note_sequence = [child.tag for child in fiscal_note_root.iter()] match_results = [] for code, layout_seq in layouts_dict.items(): layout_tags = set(layout_seq) composite_score = calculate_composite_score(layout_tags, note_tags, layout_seq, note_sequence, sim_weight, order_weight) match_results.append({ 'code': code, 'similarity_score': round(calculate_similarity(layout_tags, note_tags), 3), 'order_match_score': round(calculate_order_match_score(layout_seq, note_sequence), 3), 'composite_score': round(composite_score, 3) }) sorted_results = sorted(match_results, key=lambda x: x['composite_score'], reverse=True) return sorted_results if sorted_results else 'No matching layout found' # 初始化布局字典 layouts_tree = ET.parse('layouts.xml') layouts_root = layouts_tree.getroot() layouts_dict = {} for layout in layouts_root: layout_code = layout.attrib['code'] tag_sequence = [child.tag for child in layout.iter()] layouts_dict[layout_code] = tag_sequence # 执行校验并输出前5个最优匹配 result = validate_fiscal_note('Nota 2.xml') for item in result[:5]: print(item)
关键说明
- 兼容场景支持:LCS算法自动忽略票据多出的标签,同时允许票据缺失布局中的标签,仅计算共有标签的顺序匹配度
- 权重灵活调整:如果业务更看重标签顺序,可提高
order_weight(比如设为0.7),反之则提高相似性权重 - 节点层级控制:代码默认校验所有子节点的标签顺序,若只需校验根节点的直接子节点,将
iter()替换为直接遍历节点即可
内容的提问来源于stack exchange,提问作者Jardel Galvão Rodrigues
相关产品推荐
相关产品推荐

