You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何去除目录列表中重复文件名并在XML标签中仅写入一次

问题:修复重复添加PDF通知标签的Python脚本

问题背景

文件树结构如下:

2_Product
    2-1_CategoryName1_Product
        2-1-1_Name1_Product
            LLL_nomenclature1_product.zip
                LLL_nomenclature1_product (folder)
                    notice_nomenclature1.pdf
            LLL_nomenclature1_product_metadata.xml
            LLL_nomenclature2_product.zip
                LLL_nomenclature2_product (folder)
                    notice_nomenclature2.pdf
            LLL_nomenclature2_product_metadata.xml
            LLL_nomenclature3_product.zip
                LLL_nomenclature3_subproduct1 (folder)
                    notice_nomenclature3.pdf
                LLL_nomenclature3_subproduct2 (folder)
                    notice_nomenclature3.pdf
                LLL_nomenclature3_subproduct3 (folder)
                    notice_nomenclature3.pdf
            LLL_nomenclature3_product_metadata.xml
            ... etc
        2-1-2_Name2_Product
        2-1-3_ ...etc
    2-2_CategoryName2_Product
        2-2-1_ ...
        2-2-2_ ...
    ... etc

编写的Python脚本用于在压缩包中搜索notice_nomenclatureX.pdf,并在对应XML文件中添加<notice_pdf>标签,原脚本代码:

import os
import xml.etree.ElementTree as ET
import zipfile

for root, dirs, files in os.walk("."):
    for folder_ext in files:
        if folder_ext[-4:] == '.zip' and folder_ext[:3] == 'LLL': 
            filePath3 = os.path.join(root, folder_ext)
            zip_folder = zipfile.ZipFile(filePath3)
            zipfile_paths = zip_folder.namelist() 
            for paths in zipfile_paths: 
                zipfiles = os.path.basename(paths)
                if zipfiles[-4:] == '.pdf' and zipfiles[:3] == 'not':
                    notice_name = zipfiles 
                    for prdt in files: 
                        if prdt[-4:] == '.xml' and prdt[:-13] == folder_ext[:-4] : 
                            filePath4 = os.path.join(root, prdt) 
                            xml_produit = ET.parse(filePath4) 
                            root_produit = xml_produit.getroot() 
                            notice_tag = ET.SubElement(root_produit, "notice_pdf") 
                            notice_tag.text = notice_name 
                            ET.indent(root_produit) 
                            xml_produit.write(filePath4, encoding='utf-8', xml_declaration=True, method='xml', short_empty_elements=False)

问题现象

  • 处理nomenclature1和nomenclature2时正常,XML仅添加1个对应标签
  • 处理nomenclature3时,因压缩包内有3个同名PDF,XML中重复添加3次相同标签,不符合需求

尝试的方法及错误

尝试用列表去重,但错误地将notice_name设为列表,引发类型错误:

...
new_list = []
for paths in zipfile_paths:
    zipfiles = os.path.basename(paths)
    if zipfiles[-4:] == '.pdf' and zipfiles[:3] == 'not': 
        if zipfiles not in new_list:
            new_list.append(zipfiles)
            notice_name = new_list
... etc

报错信息:

TypeError: write() argument must be str, not list

解决方案

核心思路:先提取压缩包内所有唯一的通知PDF名称,再针对每个唯一名称仅向XML添加一次标签(同时检查XML中是否已存在该标签,避免重复处理)。

修改后的完整脚本:

import os
import xml.etree.ElementTree as ET
import zipfile

for root, dirs, files in os.walk("."):
    for folder_ext in files:
        if folder_ext.endswith('.zip') and folder_ext.startswith('LLL'): 
            zip_path = os.path.join(root, folder_ext)
            # 直接拼接对应XML文件名,避免遍历所有文件
            xml_filename = f"{folder_ext[:-4]}_metadata.xml"
            xml_path = os.path.join(root, xml_filename)
            
            # 跳过不存在的XML文件
            if not os.path.exists(xml_path):
                continue
                
            # 用集合自动去重,提取压缩包内唯一的通知PDF名称
            unique_notices = set()
            with zipfile.ZipFile(zip_path, 'r') as zip_folder:
                for path in zip_folder.namelist():
                    filename = os.path.basename(path)
                    if filename.endswith('.pdf') and filename.startswith('notice'):
                        unique_notices.add(filename)
            
            # 处理XML文件,添加唯一标签
            xml_tree = ET.parse(xml_path)
            xml_root = xml_tree.getroot()
            
            # 检查已有标签,避免重复添加(可选但推荐)
            existing_notices = [elem.text for elem in xml_root.findall('notice_pdf')]
            
            for notice_name in unique_notices:
                if notice_name not in existing_notices:
                    notice_tag = ET.SubElement(xml_root, "notice_pdf")
                    notice_tag.text = notice_name
            
            ET.indent(xml_root)
            xml_tree.write(xml_path, encoding='utf-8', xml_declaration=True, method='xml', short_empty_elements=False)

关键修改点

  1. 集合去重:使用set()自动处理压缩包内的重复PDF名称,无需手动判断
  2. 优化XML定位:通过zip文件名直接拼接对应XML路径,减少嵌套循环,提升效率
  3. 重复标签检查:提前查询XML中已有的notice_pdf标签内容,避免多次运行脚本时重复添加
  4. 资源管理优化:用with语句自动管理zip文件的打开与关闭,避免资源泄漏

内容的提问来源于stack exchange,提问作者Camille

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 13:37:05