You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历多子文件夹时统计文件总数及代码计数异常问题求助

解决方案

先看修改后的完整代码,解决了总文件数统计、重复循环和潜在报错问题:

def duplicatechecker():
    logger.info("Duplicate Checker: START")
    logger.info("The process to check for duplicate files has been initiated.")
    
    total_files = 0  # 初始化总文件计数器
    invoice_number = {}  # 移到外层可跨文件夹检测重复,若需单文件夹独立检测则移回内部
    
    for dir_path in dirs:  # 直接遍历路径,避免索引冗余
        DATA_DIR = Path(dir_path)
        files = sorted(DATA_DIR.glob('*.xml'))
        dir_file_count = len(files)
        total_files += dir_file_count  # 累加当前文件夹文件数到总数
        
        # 修复f-string与format混用的错误写法
        logger.info(f"The number of files in the directory {dir_path} is {dir_file_count}")
        
        duplicateFiles = []
        client_id = DATA_DIR.name
        
        # 直接遍历文件对象,无需索引循环
        for file in files:
            tree = ET.parse(file)
            root = tree.getroot()
            record = root.findall('record')
            
            # 直接遍历record元素
            for rec in record:
                inv = rec.find('invoice_number').text
                if inv in invoice_number:
                    duplicateFiles.append((inv, file))
                    duplicateFiles.append((inv, invoice_number[inv]))
                    print(f"Duplicate found: invoice number {inv}, file {file}, other file {invoice_number[inv]}")
                    copy_1 = os.path.basename(str(file))
                    copy_2 = os.path.basename(str(invoice_number[inv]))
                    logger.info(f"{client_id} : Duplicate : {inv} : {copy_1[:-3]}pdf : {copy_2[:-3]}pdf")
                else:
                    invoice_number[inv] = file
        
        logger.info(f"The number of duplicate files found in {dir_path} is {len(duplicateFiles)//2}")
        # 避免空文件夹时索引报错,改用文件夹路径输出
        logger.info(f"Duplicate Checker: END for directory {dir_path}")
    
    # 最终输出所有文件夹的总文件数
    logger.info(f"Total number of files across all directories: {total_files}")
    logger.info("Duplicate Checker: COMPLETE")

duplicatechecker()
issuechecker()
logger.info("The process has been completed.")
logger.info("data_capture_issue uploaded")

关键修改说明

  • 总文件数统计:新增total_files变量,遍历每个文件夹时累加当前文件夹的文件数量,最后统一输出总数。
  • 修复重复循环问题:若第二个子文件夹重复处理,大概率是dirs列表中存在重复路径,可先通过dirs = list(set(dirs))去重(路径为字符串时可用,Path对象需先转字符串)。同时改用直接遍历路径的写法,减少索引相关错误。
  • 移除冗余格式化:原代码f"...".format(len(files))是错误用法,f-string已完成变量替换,无需额外调用format方法。
  • 避免索引报错:原代码末尾files[i]会在空文件夹时抛出IndexError,改为输出文件夹路径更安全。
  • 优化循环逻辑:将所有for i in range(len(xxx))的索引遍历改为直接遍历元素,代码更简洁易读,降低出错概率。

若不需要跨文件夹检测重复发票号,只需每个文件夹独立检测,把invoice_number = {}移回每个文件夹循环的内部即可。

内容的提问来源于stack exchange,提问作者Flint_Lockwood

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 04:40:34