You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python校验XML票据与XML布局的标签顺序及兼容性?

问题

作为开发新手,我正在编写Python脚本校验XML票据(Fiscal Note)与XML布局(Layout)的兼容性。目前已用xml.etree.ElementTree实现了标签相似性校验,现在需要升级脚本,实现标签顺序校验,同时兼容以下场景:

  • 票据缺少布局中的某些标签
  • 票据多出布局中没有的标签

现有代码如下:

import xml.etree.ElementTree as ET

# Function to calculate similarity between two sets
def calculate_similarity(set1, set2):
    intersection = set1.intersection(set2)
    smallest_set_size = min(len(set1), len(set2))
    if smallest_set_size == 0:
        return 0
    similarity = len(intersection) / smallest_set_size
    return similarity


# Function to validate a fiscal note against the layouts
def validate_fiscal_note(fiscal_note_xml, similarity_threshold=0.1):
    fiscal_note_tree = ET.parse(fiscal_note_xml)
    fiscal_note_root = fiscal_note_tree.getroot()
    
    fiscal_note_tags = {child.tag for child in fiscal_note_root.iter()}

    best_match_code = None
    best_match_similarity = 0
    best_match_code_list = []

    for code, layout_tags in layouts_dict.items():
        similarity = calculate_similarity(layout_tags, fiscal_note_tags)
        best_match_code_list.append({'code': code, 'similarity': similarity})
    
    sorted_list = sorted(best_match_code_list, key=lambda x: x['similarity'])

    if not sorted_list:
        return 'Any layout was finded'
    else:
        return sorted_list

# Parse the layouts XML
layouts_tree = ET.parse('layouts.xml')
layouts_root = layouts_tree.getroot()

# Build a dictionary of layouts
layouts_dict = {}
for layout in layouts_root:
    layout_code = layout.attrib['code']  # Make sure this matches your XML structure; might need to adjust
    tags = {child.tag for child in layout}
    layouts_dict[layout_code] = tags

# Validate the fiscal note and print the result
result = validate_fiscal_note('Nota 2.xml')
print(result[-5:])

解决方案

1. 重构布局字典:存储标签顺序而非集合

原代码用集合存储标签丢失了顺序信息,需要改为存储标签顺序列表:

# 解析布局XML并构建字典(存储标签顺序)
layouts_tree = ET.parse('layouts.xml')
layouts_root = layouts_tree.getroot()

layouts_dict = {}
for layout in layouts_root:
    layout_code = layout.attrib['code']
    # 按XML中的实际顺序收集所有标签(含子节点),若只需直接子节点则替换为 `layout`
    tag_sequence = [child.tag for child in layout.iter()]
    layouts_dict[layout_code] = tag_sequence

2. 实现顺序匹配算法:最长公共子序列(LCS)

LCS算法可以兼容票据缺标签、多标签的场景,只关注两者共有标签的顺序匹配度:

def calculate_order_match_score(layout_sequence, note_sequence):
    m, n = len(layout_sequence), len(note_sequence)
    # 构建动态规划表
    dp = [[0]*(n+1) for _ in range(m+1)]
    
    for i in range(1, m+1):
        for j in range(1, n+1):
            if layout_sequence[i-1] == note_sequence[j-1]:
                dp[i][j] = dp[i-1][j-1] + 1
            else:
                dp[i][j] = max(dp[i-1][j], dp[i][j-1])
    
    # 顺序得分 = LCS长度 / 布局标签总数(保证布局顺序的优先级)
    return dp[m][n] / m if m != 0 else 0.0

3. 整合相似性与顺序得分,计算综合匹配度

可根据业务需求调整两者的权重(比如各占50%):

def calculate_composite_score(layout_tags, note_tags, layout_seq, note_seq, sim_weight=0.5, order_weight=0.5):
    sim_score = calculate_similarity(layout_tags, note_tags)
    order_score = calculate_order_match_score(layout_seq, note_seq)
    return sim_weight * sim_score + order_weight * order_score

4. 重构校验函数,同时输出多维度得分

更新校验逻辑,同时计算相似性、顺序匹配度和综合得分:

def validate_fiscal_note(fiscal_note_xml, sim_weight=0.5, order_weight=0.5):
    fiscal_note_tree = ET.parse(fiscal_note_xml)
    fiscal_note_root = fiscal_note_tree.getroot()
    
    # 收集票据的标签集合和顺序序列(若只需直接子节点则替换为 `fiscal_note_root`)
    note_tags = {child.tag for child in fiscal_note_root.iter()}
    note_sequence = [child.tag for child in fiscal_note_root.iter()]
    
    match_results = []
    for code, layout_seq in layouts_dict.items():
        layout_tags = set(layout_seq)
        sim_score = calculate_similarity(layout_tags, note_tags)
        order_score = calculate_order_match_score(layout_seq, note_sequence)
        composite_score = calculate_composite_score(layout_tags, note_tags, layout_seq, note_sequence, sim_weight, order_weight)
        
        match_results.append({
            'code': code,
            'similarity_score': round(sim_score, 3),
            'order_match_score': round(order_score, 3),
            'composite_score': round(composite_score, 3)
        })
    
    # 按综合得分降序排序
    sorted_results = sorted(match_results, key=lambda x: x['composite_score'], reverse=True)
    return sorted_results if sorted_results else 'No matching layout found'

5. 完整整合代码

import xml.etree.ElementTree as ET

def calculate_similarity(set1, set2):
    intersection = set1.intersection(set2)
    smallest_set_size = min(len(set1), len(set2))
    return len(intersection) / smallest_set_size if smallest_set_size != 0 else 0.0

def calculate_order_match_score(layout_sequence, note_sequence):
    m, n = len(layout_sequence), len(note_sequence)
    dp = [[0]*(n+1) for _ in range(m+1)]
    
    for i in range(1, m+1):
        for j in range(1, n+1):
            if layout_sequence[i-1] == note_sequence[j-1]:
                dp[i][j] = dp[i-1][j-1] + 1
            else:
                dp[i][j] = max(dp[i-1][j], dp[i][j-1])
    
    return dp[m][n] / m if m != 0 else 0.0

def calculate_composite_score(layout_tags, note_tags, layout_seq, note_seq, sim_weight=0.5, order_weight=0.5):
    sim_score = calculate_similarity(layout_tags, note_tags)
    order_score = calculate_order_match_score(layout_seq, note_seq)
    return sim_weight * sim_score + order_weight * order_score

def validate_fiscal_note(fiscal_note_xml, sim_weight=0.5, order_weight=0.5):
    fiscal_note_tree = ET.parse(fiscal_note_xml)
    fiscal_note_root = fiscal_note_tree.getroot()
    
    # 若只需校验根节点直接子节点,替换为:
    # note_tags = {child.tag for child in fiscal_note_root}
    # note_sequence = [child.tag for child in fiscal_note_root]
    note_tags = {child.tag for child in fiscal_note_root.iter()}
    note_sequence = [child.tag for child in fiscal_note_root.iter()]
    
    match_results = []
    for code, layout_seq in layouts_dict.items():
        layout_tags = set(layout_seq)
        composite_score = calculate_composite_score(layout_tags, note_tags, layout_seq, note_sequence, sim_weight, order_weight)
        
        match_results.append({
            'code': code,
            'similarity_score': round(calculate_similarity(layout_tags, note_tags), 3),
            'order_match_score': round(calculate_order_match_score(layout_seq, note_sequence), 3),
            'composite_score': round(composite_score, 3)
        })
    
    sorted_results = sorted(match_results, key=lambda x: x['composite_score'], reverse=True)
    return sorted_results if sorted_results else 'No matching layout found'

# 初始化布局字典
layouts_tree = ET.parse('layouts.xml')
layouts_root = layouts_tree.getroot()
layouts_dict = {}
for layout in layouts_root:
    layout_code = layout.attrib['code']
    tag_sequence = [child.tag for child in layout.iter()]
    layouts_dict[layout_code] = tag_sequence

# 执行校验并输出前5个最优匹配
result = validate_fiscal_note('Nota 2.xml')
for item in result[:5]:
    print(item)

关键说明

  • 兼容场景支持:LCS算法自动忽略票据多出的标签,同时允许票据缺失布局中的标签,仅计算共有标签的顺序匹配度
  • 权重灵活调整:如果业务更看重标签顺序,可提高order_weight(比如设为0.7),反之则提高相似性权重
  • 节点层级控制:代码默认校验所有子节点的标签顺序,若只需校验根节点的直接子节点,将iter()替换为直接遍历节点即可

内容的提问来源于stack exchange,提问作者Jardel Galvão Rodrigues

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 00:45:10