如何在Python中高效持续保存包含DataFrame的列表?
Hey there! Let's break down the best ways to handle your scenario—where you're generating DataFrames in a loop all day, need to append them continuously, and prioritize storage efficiency (small file sizes + fast writes).
Top Solutions for Your Use Case
1. HDF5 Format (Best for Continuous Incremental Writes)
HDF5 is perfect here because it supports direct incremental appending without needing to load all existing data into memory (critical for long-running programs). It also offers great compression that balances speed and file size.
Here's how to implement it with pandas:
import pandas as pd # Initialize the HDFStore with high compression (blosc is fast and efficient) with pd.HDFStore('continuous_data.h5', mode='a', complevel=9, complib='blosc') as store: for i in x: # Your time-consuming computation to generate a DataFrame df = your_computation_logic(i) # Append the new DataFrame to the HDF file # Use a unique key (like the iteration index) to avoid conflicts store.append(f'dataset_{i}', df, format='table')
Pros:
- No need to load existing data to append—writes directly to disk
- Supports compression (blosc is ideal for speed + compression ratio)
- Lets you read specific DataFrames later by key, or query across all data
- Handles large datasets without blowing up memory usage
Cons:
- Tied to Python/pandas ecosystem (less cross-language support)
- Requires the
pytableslibrary (usually installed automatically with pandas)
2. Parquet Format (Best for Extreme Compression & Long-Term Storage)
Parquet is a columnar storage format that delivers industry-leading compression ratios (great for large files) and works across many programming languages. While native single-file append is limited, you can use a multi-file approach or leverage pyarrow's append capabilities.
Option A: Multi-File Parquet (Simpler, Scalable)
Generate a small Parquet file per iteration, then combine them later if needed:
import pandas as pd for i in x: df = your_computation_logic(i) # Save each DataFrame as a separate Parquet file with a unique name df.to_parquet(f'data_chunk_{i}.parquet', compression='snappy') # Later, to read all chunks into a single DataFrame: import glob all_chunks = pd.concat([pd.read_parquet(file) for file in glob.glob('data_chunk_*.parquet')], ignore_index=True)
Option B: Single-File Append (Requires PyArrow)
If you prefer a single file, use pyarrow's append mode (requires pyarrow >= 0.17.0):
import pandas as pd import pyarrow as pa import pyarrow.parquet as pq first_iteration = True for i in x: df = your_computation_logic(i) table = pa.Table.from_pandas(df) if first_iteration: # Create the initial file pq.write_table(table, 'continuous_data.parquet', compression='snappy') first_iteration = False else: # Append to the existing file pq.write_to_dataset(table, root_path='continuous_data.parquet', append=True)
Pros:
- Unbeatable compression (smallest file sizes for large datasets)
- Cross-language support (works with Spark, R, Java, etc.)
- Fast read speeds for analytical queries
Cons:
- Single-file append has version dependencies
- Multi-file approach requires cleanup/merging later
3. Feather Format (Best for Blazing-Fast Reads/Writes)
Feather is designed for ultra-fast data exchange between memory and disk. It's great if you need to frequently read/write the data during your program, though full-file appending (reading all data, merging, re-writing) can get slow with very large datasets.
import pandas as pd import pyarrow.feather as feather first_iteration = True for i in x: df = your_computation_logic(i) if first_iteration: feather.write_feather(df, 'continuous_data.feather') first_iteration = False else: # Note: This method reads existing data into memory first existing_data = feather.read_feather('continuous_data.feather') combined_data = pd.concat([existing_data, df], ignore_index=True) feather.write_feather(combined_data, 'continuous_data.feather')
Pros:
- Fastest read/write speeds of all formats
- Lightweight, easy to use
Cons:
- Full-file append requires loading all data into memory (not ideal for huge datasets)
- Less compression than Parquet/HDF5
Final Recommendations
- If you need continuous, memory-efficient appending and plan to work with the data in Python: Go with HDF5.
- If you prioritize smallest file sizes and cross-language compatibility: Choose Parquet (multi-file is safest for long-running loops).
- If speed is your top priority and dataset size is manageable: Use Feather.
内容的提问来源于stack exchange,提问作者Eghbal

