You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中高效持续保存包含DataFrame的列表?

Efficiently Saving a Growing List of DataFrames in Python (Continuous Appending)

Hey there! Let's break down the best ways to handle your scenario—where you're generating DataFrames in a loop all day, need to append them continuously, and prioritize storage efficiency (small file sizes + fast writes).

Top Solutions for Your Use Case

1. HDF5 Format (Best for Continuous Incremental Writes)

HDF5 is perfect here because it supports direct incremental appending without needing to load all existing data into memory (critical for long-running programs). It also offers great compression that balances speed and file size.

Here's how to implement it with pandas:

import pandas as pd

# Initialize the HDFStore with high compression (blosc is fast and efficient)
with pd.HDFStore('continuous_data.h5', mode='a', complevel=9, complib='blosc') as store:
    for i in x:
        # Your time-consuming computation to generate a DataFrame
        df = your_computation_logic(i)
        
        # Append the new DataFrame to the HDF file
        # Use a unique key (like the iteration index) to avoid conflicts
        store.append(f'dataset_{i}', df, format='table')

Pros:

  • No need to load existing data to append—writes directly to disk
  • Supports compression (blosc is ideal for speed + compression ratio)
  • Lets you read specific DataFrames later by key, or query across all data
  • Handles large datasets without blowing up memory usage

Cons:

  • Tied to Python/pandas ecosystem (less cross-language support)
  • Requires the pytables library (usually installed automatically with pandas)

2. Parquet Format (Best for Extreme Compression & Long-Term Storage)

Parquet is a columnar storage format that delivers industry-leading compression ratios (great for large files) and works across many programming languages. While native single-file append is limited, you can use a multi-file approach or leverage pyarrow's append capabilities.

Option A: Multi-File Parquet (Simpler, Scalable)

Generate a small Parquet file per iteration, then combine them later if needed:

import pandas as pd

for i in x:
    df = your_computation_logic(i)
    # Save each DataFrame as a separate Parquet file with a unique name
    df.to_parquet(f'data_chunk_{i}.parquet', compression='snappy')

# Later, to read all chunks into a single DataFrame:
import glob
all_chunks = pd.concat([pd.read_parquet(file) for file in glob.glob('data_chunk_*.parquet')], ignore_index=True)

Option B: Single-File Append (Requires PyArrow)

If you prefer a single file, use pyarrow's append mode (requires pyarrow >= 0.17.0):

import pandas as pd
import pyarrow as pa
import pyarrow.parquet as pq

first_iteration = True
for i in x:
    df = your_computation_logic(i)
    table = pa.Table.from_pandas(df)
    
    if first_iteration:
        # Create the initial file
        pq.write_table(table, 'continuous_data.parquet', compression='snappy')
        first_iteration = False
    else:
        # Append to the existing file
        pq.write_to_dataset(table, root_path='continuous_data.parquet', append=True)

Pros:

  • Unbeatable compression (smallest file sizes for large datasets)
  • Cross-language support (works with Spark, R, Java, etc.)
  • Fast read speeds for analytical queries

Cons:

  • Single-file append has version dependencies
  • Multi-file approach requires cleanup/merging later

3. Feather Format (Best for Blazing-Fast Reads/Writes)

Feather is designed for ultra-fast data exchange between memory and disk. It's great if you need to frequently read/write the data during your program, though full-file appending (reading all data, merging, re-writing) can get slow with very large datasets.

import pandas as pd
import pyarrow.feather as feather

first_iteration = True
for i in x:
    df = your_computation_logic(i)
    
    if first_iteration:
        feather.write_feather(df, 'continuous_data.feather')
        first_iteration = False
    else:
        # Note: This method reads existing data into memory first
        existing_data = feather.read_feather('continuous_data.feather')
        combined_data = pd.concat([existing_data, df], ignore_index=True)
        feather.write_feather(combined_data, 'continuous_data.feather')

Pros:

  • Fastest read/write speeds of all formats
  • Lightweight, easy to use

Cons:

  • Full-file append requires loading all data into memory (not ideal for huge datasets)
  • Less compression than Parquet/HDF5

Final Recommendations

  • If you need continuous, memory-efficient appending and plan to work with the data in Python: Go with HDF5.
  • If you prioritize smallest file sizes and cross-language compatibility: Choose Parquet (multi-file is safest for long-running loops).
  • If speed is your top priority and dataset size is manageable: Use Feather.

内容的提问来源于stack exchange,提问作者Eghbal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:05:25