You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas筛选文件夹最早修改日期CSV时无法获取正确修改时间问题

问题根因与解决方案

1 现有代码的直接逻辑错误

你当前的循环写法存在赋值逻辑问题:每次遍历单个文件时,直接对output['Creation date']、output['Modif. date']整列赋值,最终所有行的日期都会被覆盖为最后一个遍历文件的日期,这是你得到的所有日期完全一致的核心原因。

2 修复后的基础实现代码

import os
import glob
import time
import pandas as pd

# 配置参数
path_input = r"S:\Production\Verspaning\Grootverspaning\Meetr..."
format_str = "*.csv"

# 读取所有csv文件路径
csv_files = glob.glob(os.path.join(path_input, format_str))
output = pd.DataFrame(csv_files, columns=["file path"])

# 逐行赋值日期属性
for idx, row in output.iterrows():
    file_path = row["file path"]
    output.loc[idx, "Creation date"] = time.ctime(os.path.getctime(file_path))
    output.loc[idx, "Modif. date"] = time.ctime(os.path.getmtime(file_path))

3 共享盘日期读取兼容方案

如果修复上述代码后,读取的日期仍和Windows资源管理器显示的不一致,是因为S盘为SMB网络共享盘,跨平台的os模块接口无法正确读取Windows NTFS分区的原生属性,需要使用pywin32库调用Windows原生API获取属性:

  • 先安装依赖:pip install pywin32
  • 替换日期读取逻辑:
import win32file
import win32con

def get_windows_file_date(file_path):
    # 读取Windows原生文件时间属性
    handle = win32file.CreateFile(
        file_path, win32con.GENERIC_READ, 
        win32con.FILE_SHARE_READ | win32con.FILE_SHARE_WRITE,
        None, win32con.OPEN_EXISTING, 0, None
    )
    create_time, access_time, modify_time = win32file.GetFileTime(handle)
    win32file.CloseHandle(handle)
    # 转换为可读格式
    return time.ctime(create_time), time.ctime(modify_time)

# 调用方法获取日期
for idx, row in output.iterrows():
    c_time, m_time = get_windows_file_date(row["file path"])
    output.loc[idx, "Creation date"] = c_time
    output.loc[idx, "Modif. date"] = m_time

4 重复文件去重逻辑

拿到正确的修改日期后,可按文件内容哈希(或文件名+文件大小的组合特征)分组,每组保留修改日期最早的文件即可:

# 示例:按文件内容MD5去重,可根据实际场景替换为你的重复判断规则
import hashlib

def get_file_md5(file_path):
    md5_hash = hashlib.md5()
    with open(file_path, "rb") as f:
        # 大文件分块读取避免内存溢出
        for chunk in iter(lambda: f.read(4096), b""):
            md5_hash.update(chunk)
    return md5_hash.hexdigest()

# 新增文件哈希列
output["file_md5"] = output["file path"].apply(get_file_md5)
# 按哈希分组,保留修改日期最早的记录
output["Modif. timestamp"] = output["Modif. date"].apply(lambda x: time.mktime(time.strptime(x)))
output = output.sort_values("Modif. timestamp").drop_duplicates("file_md5", keep="first")

内容的提问来源于stack exchange,提问作者GuillermoDC

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 03:06:03