如何去除目录列表中重复文件名并在XML标签中仅写入一次
问题:修复重复添加PDF通知标签的Python脚本
问题背景
文件树结构如下:
2_Product 2-1_CategoryName1_Product 2-1-1_Name1_Product LLL_nomenclature1_product.zip LLL_nomenclature1_product (folder) notice_nomenclature1.pdf LLL_nomenclature1_product_metadata.xml LLL_nomenclature2_product.zip LLL_nomenclature2_product (folder) notice_nomenclature2.pdf LLL_nomenclature2_product_metadata.xml LLL_nomenclature3_product.zip LLL_nomenclature3_subproduct1 (folder) notice_nomenclature3.pdf LLL_nomenclature3_subproduct2 (folder) notice_nomenclature3.pdf LLL_nomenclature3_subproduct3 (folder) notice_nomenclature3.pdf LLL_nomenclature3_product_metadata.xml ... etc 2-1-2_Name2_Product 2-1-3_ ...etc 2-2_CategoryName2_Product 2-2-1_ ... 2-2-2_ ... ... etc
编写的Python脚本用于在压缩包中搜索notice_nomenclatureX.pdf,并在对应XML文件中添加<notice_pdf>标签,原脚本代码:
import os import xml.etree.ElementTree as ET import zipfile for root, dirs, files in os.walk("."): for folder_ext in files: if folder_ext[-4:] == '.zip' and folder_ext[:3] == 'LLL': filePath3 = os.path.join(root, folder_ext) zip_folder = zipfile.ZipFile(filePath3) zipfile_paths = zip_folder.namelist() for paths in zipfile_paths: zipfiles = os.path.basename(paths) if zipfiles[-4:] == '.pdf' and zipfiles[:3] == 'not': notice_name = zipfiles for prdt in files: if prdt[-4:] == '.xml' and prdt[:-13] == folder_ext[:-4] : filePath4 = os.path.join(root, prdt) xml_produit = ET.parse(filePath4) root_produit = xml_produit.getroot() notice_tag = ET.SubElement(root_produit, "notice_pdf") notice_tag.text = notice_name ET.indent(root_produit) xml_produit.write(filePath4, encoding='utf-8', xml_declaration=True, method='xml', short_empty_elements=False)
问题现象
- 处理
nomenclature1和nomenclature2时正常,XML仅添加1个对应标签 - 处理
nomenclature3时,因压缩包内有3个同名PDF,XML中重复添加3次相同标签,不符合需求
尝试的方法及错误
尝试用列表去重,但错误地将notice_name设为列表,引发类型错误:
... new_list = [] for paths in zipfile_paths: zipfiles = os.path.basename(paths) if zipfiles[-4:] == '.pdf' and zipfiles[:3] == 'not': if zipfiles not in new_list: new_list.append(zipfiles) notice_name = new_list ... etc
报错信息:
TypeError: write() argument must be str, not list
解决方案
核心思路:先提取压缩包内所有唯一的通知PDF名称,再针对每个唯一名称仅向XML添加一次标签(同时检查XML中是否已存在该标签,避免重复处理)。
修改后的完整脚本:
import os import xml.etree.ElementTree as ET import zipfile for root, dirs, files in os.walk("."): for folder_ext in files: if folder_ext.endswith('.zip') and folder_ext.startswith('LLL'): zip_path = os.path.join(root, folder_ext) # 直接拼接对应XML文件名,避免遍历所有文件 xml_filename = f"{folder_ext[:-4]}_metadata.xml" xml_path = os.path.join(root, xml_filename) # 跳过不存在的XML文件 if not os.path.exists(xml_path): continue # 用集合自动去重,提取压缩包内唯一的通知PDF名称 unique_notices = set() with zipfile.ZipFile(zip_path, 'r') as zip_folder: for path in zip_folder.namelist(): filename = os.path.basename(path) if filename.endswith('.pdf') and filename.startswith('notice'): unique_notices.add(filename) # 处理XML文件,添加唯一标签 xml_tree = ET.parse(xml_path) xml_root = xml_tree.getroot() # 检查已有标签,避免重复添加(可选但推荐) existing_notices = [elem.text for elem in xml_root.findall('notice_pdf')] for notice_name in unique_notices: if notice_name not in existing_notices: notice_tag = ET.SubElement(xml_root, "notice_pdf") notice_tag.text = notice_name ET.indent(xml_root) xml_tree.write(xml_path, encoding='utf-8', xml_declaration=True, method='xml', short_empty_elements=False)
关键修改点
- 集合去重:使用
set()自动处理压缩包内的重复PDF名称,无需手动判断 - 优化XML定位:通过zip文件名直接拼接对应XML路径,减少嵌套循环,提升效率
- 重复标签检查:提前查询XML中已有的
notice_pdf标签内容,避免多次运行脚本时重复添加 - 资源管理优化:用
with语句自动管理zip文件的打开与关闭,避免资源泄漏
内容的提问来源于stack exchange,提问作者Camille
相关产品推荐
相关产品推荐

