You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Jupyter Notebook中自动缓存pandas读取的文件?

文件读取自动缓存最优实现方案

方案1:轻量自定义装饰器(无额外依赖)

这是最通用的实现方式,无需安装第三方库,逻辑可自主扩展:

  • 核心逻辑:用全局字典存储缓存映射,key为文件绝对路径,value存储「文件最后修改时间、读取结果」
  • 每次调用读取函数时,先获取目标文件的最新修改时间,和缓存记录比对:时间一致则直接返回缓存数据,否则重新读取并更新缓存
  • 实现代码如下:
import os
import pandas as pd
from functools import wraps

# 全局缓存存储:key为文件绝对路径,value为(最后修改时间, 数据内容)
file_cache = {}

def cache_file_read(read_func):
    @wraps(read_func)
    def wrapper(file_path, *args, **kwargs):
        # 转绝对路径避免相对路径重复缓存,有软链接场景可替换为os.path.realpath
        abs_path = os.path.abspath(file_path)
        # 获取文件最新修改时间
        current_mtime = os.path.getmtime(abs_path)
        
        # 校验缓存有效性
        if abs_path in file_cache:
            cached_mtime, cached_data = file_cache[abs_path]
            if current_mtime == cached_mtime:
                return cached_data
        
        # 缓存失效,重新读取文件
        data = read_func(file_path, *args, **kwargs)
        file_cache[abs_path] = (current_mtime, data)
        return data
    return wrapper
  • 使用方式:
# 给需要的读取函数添加装饰器,read_csv、read_excel等都可以适配
@cache_file_read
def cached_read_table(file_path, *args, **kwargs):
    return pd.read_table(file_path, *args, **kwargs)

# 调用方式和原生接口一致,自动适配缓存逻辑
data = cached_read_table('my_data.txt')

方案2:基于joblib实现持久化缓存(适合重启场景)

如果需要重启Python进程/Notebook内核后依然保留缓存,可以用joblib的持久化缓存能力,需要先安装依赖:pip install joblib

  • 核心逻辑:把文件修改时间作为缓存key的一部分,自动触发缓存更新,缓存会落地到本地磁盘
  • 实现代码如下:
import os
import pandas as pd
from joblib import Memory

# 初始化缓存存储目录,verbose=0关闭日志输出
memory = Memory(location='./pandas_read_cache', verbose=0)

def cached_read_table(file_path, *args, **kwargs):
    current_mtime = os.path.getmtime(file_path)
    # 将修改时间作为缓存判断依据,忽略file_path参数避免路径变化触发重缓存
    @memory.cache(ignore=['file_path'])
    def _read(_mtime):
        return pd.read_table(file_path, *args, **kwargs)
    return _read(current_mtime)

# 调用方式不变
data = cached_read_table('my_data.txt')

优化建议

  • 内存占用控制:如果读取的文件体积较大,可以把全局缓存替换为LRU淘汰策略的缓存,使用functools.lru_cache或者cachetools库的LRUCache即可限制最大内存占用
  • 多文件适配:如果需要同时缓存多种读取函数的结果,可在缓存key中加入读取函数标识,避免不同读取逻辑混用同一个文件的缓存

内容的提问来源于stack exchange,提问作者rhombidodecahedron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 06:06:05