如何在Jupyter Notebook中自动缓存pandas读取的文件?
文件读取自动缓存最优实现方案
方案1:轻量自定义装饰器(无额外依赖)
这是最通用的实现方式,无需安装第三方库,逻辑可自主扩展:
- 核心逻辑:用全局字典存储缓存映射,key为文件绝对路径,value存储「文件最后修改时间、读取结果」
- 每次调用读取函数时,先获取目标文件的最新修改时间,和缓存记录比对:时间一致则直接返回缓存数据,否则重新读取并更新缓存
- 实现代码如下:
import os import pandas as pd from functools import wraps # 全局缓存存储:key为文件绝对路径,value为(最后修改时间, 数据内容) file_cache = {} def cache_file_read(read_func): @wraps(read_func) def wrapper(file_path, *args, **kwargs): # 转绝对路径避免相对路径重复缓存,有软链接场景可替换为os.path.realpath abs_path = os.path.abspath(file_path) # 获取文件最新修改时间 current_mtime = os.path.getmtime(abs_path) # 校验缓存有效性 if abs_path in file_cache: cached_mtime, cached_data = file_cache[abs_path] if current_mtime == cached_mtime: return cached_data # 缓存失效,重新读取文件 data = read_func(file_path, *args, **kwargs) file_cache[abs_path] = (current_mtime, data) return data return wrapper
- 使用方式:
# 给需要的读取函数添加装饰器,read_csv、read_excel等都可以适配 @cache_file_read def cached_read_table(file_path, *args, **kwargs): return pd.read_table(file_path, *args, **kwargs) # 调用方式和原生接口一致,自动适配缓存逻辑 data = cached_read_table('my_data.txt')
方案2:基于joblib实现持久化缓存(适合重启场景)
如果需要重启Python进程/Notebook内核后依然保留缓存,可以用joblib的持久化缓存能力,需要先安装依赖:pip install joblib
- 核心逻辑:把文件修改时间作为缓存key的一部分,自动触发缓存更新,缓存会落地到本地磁盘
- 实现代码如下:
import os import pandas as pd from joblib import Memory # 初始化缓存存储目录,verbose=0关闭日志输出 memory = Memory(location='./pandas_read_cache', verbose=0) def cached_read_table(file_path, *args, **kwargs): current_mtime = os.path.getmtime(file_path) # 将修改时间作为缓存判断依据,忽略file_path参数避免路径变化触发重缓存 @memory.cache(ignore=['file_path']) def _read(_mtime): return pd.read_table(file_path, *args, **kwargs) return _read(current_mtime) # 调用方式不变 data = cached_read_table('my_data.txt')
优化建议
- 内存占用控制:如果读取的文件体积较大,可以把全局缓存替换为LRU淘汰策略的缓存,使用
functools.lru_cache或者cachetools库的LRUCache即可限制最大内存占用 - 多文件适配:如果需要同时缓存多种读取函数的结果,可在缓存key中加入读取函数标识,避免不同读取逻辑混用同一个文件的缓存
内容的提问来源于stack exchange,提问作者rhombidodecahedron
相关产品推荐
相关产品推荐

