You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python项目中如何获取目录内新增的最后n个文件夹名称?

检测目录新增文件夹并批量存入SQL Server数据库

我正在开发Python项目,需要频繁循环检测指定目录里的新建文件夹,把新文件夹的名称存入SQL Server数据库。目标目录的结构是这样的:

C:\USERS\DESKTOP\RESULT
├───20220819_160033
├───20220819_162117
├───20220819_192122
├───20220820_080920
├───20220820_094355
├───20220902_081328
└───20220902_083015

目前的做法是从数据库里拿已保存的历史文件夹数量,和当前目录里的文件夹数量对比。计划是如果当前数量比历史多,算出差值n,然后把最后n个文件夹名称存进数据库,但不知道怎么实现批量获取这n个新增文件夹。

当前的代码片段:

import pandas as pd
import pyodbc, os

# Connect to SQL Server
conn = pyodbc.connect('Driver={SQL Server};'
                    'Server=server;'
                    'Database=folder;'
                    'Trusted_Connection=yes;')
cursor = conn.cursor()

cursor.execute("SELECT * FROM FileLog ORDER BY FileID DESC")
prevFileCount = cursor.fetchone()
fileCount = prevFileCount[3] # 获取历史文件夹数量

name = []
count = 0

dir_path = r'C:\USERS\DESKTOP\RESULT'
# 遍历目录统计文件夹数量
for path in os.listdir(dir_path):
    # 判断是否为文件夹
    if os.path.isdir(os.path.join(dir_path, path)):
        name.append(os.path.join(dir_path, path))
        count += 1

要解决这个问题,重点要注意这几点:

  1. 文件夹列表必须有序:os.listdir()返回的顺序是不确定的,得按文件夹名称(或者创建时间)排序——你的文件夹名称是时间戳格式,按名称排序就对应创建先后顺序了。
  2. 批量取新增文件夹:算出当前数量和历史数量的差值n后,直接从排序后的列表末尾截n个元素就行。
  3. 加上循环检测的逻辑:用定时循环实现持续监控。

完整可运行代码

import pyodbc
import os
import time

def get_db_connection():
    # 封装数据库连接,方便重复调用
    return pyodbc.connect(
        'Driver={SQL Server};'
        'Server=server;'
        'Database=folder;'
        'Trusted_Connection=yes;'
    )

def get_history_folder_count():
    # 直接查数据库里已记录的文件夹总数,比取最后一条记录更可靠
    conn = get_db_connection()
    cursor = conn.cursor()
    cursor.execute("SELECT COUNT(*) FROM FileLog")
    count = cursor.fetchone()[0]
    conn.close()
    return count

def get_current_folders(dir_path):
    # 获取目录下所有文件夹,按名称排序
    folders = []
    # 用os.scandir比os.listdir效率更高
    for entry in os.scandir(dir_path):
        if entry.is_dir():
            folders.append(entry.path)
    # 按名称排序(时间戳格式的名称会自动按时间先后排)
    folders.sort()
    return folders

def save_new_folders(new_folders):
    # 批量把新文件夹存入数据库
    if not new_folders:
        return
    conn = get_db_connection()
    cursor = conn.cursor()
    # 这里假设你的FileLog表有FolderPath字段,根据实际表结构调整
    insert_query = "INSERT INTO FileLog (FolderPath) VALUES (?)"
    # 批量插入比单条插效率高很多
    cursor.executemany(insert_query, [(path,) for path in new_folders])
    conn.commit()
    conn.close()
    print(f"已新增 {len(new_folders)} 条记录")

def monitor_folders(dir_path, interval=60):
    # 持续监控目录,interval是检测间隔(单位:秒)
    print(f"开始监控目录: {dir_path}")
    while True:
        history_count = get_history_folder_count()
        current_folders = get_current_folders(dir_path)
        current_count = len(current_folders)
        
        if current_count > history_count:
            n = current_count - history_count
            # 取最后n个就是新增的文件夹
            new_folders = current_folders[-n:]
            save_new_folders(new_folders)
        else:
            print("暂无新增文件夹")
        
        # 等指定时间后再检测
        time.sleep(interval)

if __name__ == "__main__":
    target_dir = r'C:\USERS\DESKTOP\RESULT'
    # 检测间隔设为60秒,可根据需求改
    monitor_folders(target_dir, interval=60)

额外说明

  • 如果不想按名称排序,想按实际创建时间排序,可以把folders.sort()改成:
    folders.sort(key=lambda x: os.path.getctime(x))
    
  • 原代码里通过取最后一条记录的count字段,建议换成直接查总数,避免因为数据异常导致的错误。
  • 用executemany批量插入,适合频繁检测的场景,能减少数据库连接开销。

内容的提问来源于stack exchange,提问作者adi bahadur

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 22:30:48