You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中打印重复XML文件及其对应源文件

修改后的代码实现

首先补充必要的导入语句,并调整打印逻辑以同时显示重复文件和原始文件:

from pathlib import Path
import xml.etree.ElementTree as ET

path = "path/"
def duplicatecheck():
    DATA_DIR = Path(path)
    files = sorted(DATA_DIR.glob('*.xml'))
    
    invoice_number = {}
    duplicateFiles = []

    for file in files:
        tree = ET.parse(file)
        root = tree.getroot()
        record = root.findall('record')

        for item in record:
            invoice = item.find('invoice_number').text
            if invoice in invoice_number:
                original_file = invoice_number[invoice]
                duplicateFiles.append((file, original_file))
                print(f"Duplicate file found: {file.name}, {original_file.name}")
                break
            else:
                invoice_number[invoice] = file

duplicatecheck()

关键修改点:

  • 替换索引遍历为直接遍历文件对象,代码更简洁易读
  • 检测到重复发票号时,从invoice_number字典中取出对应的原始文件
  • 打印时同时输出重复文件和原始文件的文件名(使用.name属性获取纯文件名,而非完整路径)
  • 将重复文件与原始文件的配对存入duplicateFiles列表(可选,方便后续批量处理)

输出示例:

Duplicate file found:  file (1).xml, file (a).xml
Duplicate file found:  file (2).xml, file (a).xml
Duplicate file found:  file (3).xml, file (a).xml

内容的提问来源于stack exchange,提问作者Flint_Lockwood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 20:50:35