You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速读取多个Pickle文件?Python数据加载优化求助

加速Pickle文件批量读取与合并的方法

问题背景

将我司每月的销售数据以CSV文件形式存储在文件夹中。为加快Python中的读取流程,我已将这些CSV文件转换为Pickle文件。目前使用以下代码读取所有独立Pickle文件并合并到DataFrame中:

import os, glob
import pandas as pd
import glob
import os.path

# Enter path of folder#
path = "link to the folder"

# find all pickle files
all_files = glob.glob(path + "/*.pkl")
df = pd.concat(
    (pd.read_pickle(file).assign(filename=file) for file in all_files),
    ignore_index=True,
)

共有38个独立Pickle文件,总大小95MB,但加载所有数据到DataFrame耗时56秒,求加速方法。

优化方案

1. 并行读取文件

利用多进程并行读取多个Pickle文件,避免单进程串行等待的时间浪费。可以用concurrent.futures实现:

import pandas as pd
import glob
from concurrent.futures import ProcessPoolExecutor

path = "link to the folder"
all_files = glob.glob(path + "/*.pkl")

def read_pickle_file(file):
    return pd.read_pickle(file).assign(filename=file)

with ProcessPoolExecutor() as executor:
    dfs = list(executor.map(read_pickle_file, all_files))

df = pd.concat(dfs, ignore_index=True)

注意:如果在Jupyter环境中使用,建议改用multiprocessing.Pool,避免环境兼容性问题。

2. 升级Pickle存储协议

Pickle有不同的协议版本,更高的协议(如协议4、5)读写速度更快,文件体积也更小。重新生成Pickle文件时指定高协议:

# 重新保存数据时使用协议5(Python3.8+支持)
df.to_pickle("data.pkl", protocol=5)

3. 提前合并为单个Pickle文件

既然最终要合并成一个DataFrame,不如一次性合并后保存为单个文件,后续直接读取即可,彻底避免多次小文件IO开销:

# 仅需执行一次的合并保存操作
path = "link to the folder"
all_files = glob.glob(path + "/*.pkl")
df = pd.concat((pd.read_pickle(file).assign(filename=file) for file in all_files), ignore_index=True)
df.to_pickle("combined_sales_data.pkl", protocol=5)

# 后续直接读取单个文件
df = pd.read_pickle("combined_sales_data.pkl")

4. 替换为更高效的序列化格式

用专为DataFrame设计的序列化格式替代Pickle,比如:

  • Feather:轻量级格式,读写速度远快于Pickle,还支持跨语言
  • Parquet:列式存储,压缩比高,适合大数据场景

以Feather为例:

# 一次性将CSV转为Feather格式
import pandas as pd
import glob

path = "link to the folder"
csv_files = glob.glob(path + "/*.csv")
for csv_file in csv_files:
    df = pd.read_csv(csv_file)
    df.to_feather(csv_file.replace(".csv", ".feather"))

# 读取并合并Feather文件
feather_files = glob.glob(path + "/*.feather")
df = pd.concat((pd.read_feather(file).assign(filename=file) for file in feather_files), ignore_index=True)

5. 优化存储硬件

如果文件存在机械硬盘上,IO速度会是核心瓶颈,把文件转移到固态硬盘(SSD)能大幅提升随机读写速度,直接缩短读取耗时。

内容的提问来源于stack exchange,提问作者Kavish Sewmangal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 16:24:46