You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3中使用to_hdf存储多视频DataFrame触发MemoryError

Hey there, let's tackle this memory issue you're facing with processing those temperature image CSVs into HDF files. I've dealt with similar large-scale data processing headaches before, so here are some practical, actionable solutions to try out:

1. Incremental HDF Writing (Avoid Loading All Data at Once)

Instead of loading an entire video's worth of CSVs into a single DataFrame before saving, process and append data in chunks. This keeps only a small portion of your data in memory at any time.

import pandas as pd
from pathlib import Path

def process_single_video(video_csv_dir, output_hdf_path):
    # Iterate through each CSV in the video's directory
    for csv_file in Path(video_csv_dir).glob("*.csv"):
        # Read CSV in chunks (adjust chunksize based on your memory capacity)
        for chunk in pd.read_csv(csv_file, chunksize=1000):
            # Append chunk to HDF file
            chunk.to_hdf(
                output_hdf_path,
                key="temperature_frames",
                mode="a",
                append=True,
                format="table"  # Required for appending to HDF
            )

# Example usage: process each video one by one
video_dirs = ["video_1_csvs", "video_2_csvs", ..., "video_10_csvs"]
for idx, dir_path in enumerate(video_dirs):
    output_hdf = f"video_{idx+1}_data.h5"
    process_single_video(dir_path, output_hdf)
2. Force Garbage Collection Between Videos

Python's garbage collector doesn't always free up memory immediately after variables go out of scope. Explicitly cleaning up after each video ensures you don't carry over unused data to the next processing step.

import gc

# After processing each video:
del chunk  # Delete any large variables from the current video
gc.collect()  # Trigger manual garbage collection

Add this at the end of your video processing loop to wipe the slate clean for the next video.

3. Optimize DataFrame Data Types

Temperature data rarely needs the default float64 precision. Downcasting to smaller numeric types can cut your memory usage by 50% or more.

# Specify dtypes when reading CSVs to save memory upfront
df = pd.read_csv(
    csv_file,
    dtype={
        "temperature_value": "float32",  # Use float32 instead of float64
        "frame_id": "int32"  # Use int32 instead of int64 if frame counts don't exceed 2 billion
    }
)

# Check memory usage to verify savings
print(df.info(memory_usage="deep"))
4. Use Dask for Out-of-Core Processing

If your data is way too big for even chunked Pandas, Dask is a drop-in replacement that handles datasets larger than your available RAM. It splits data into manageable partitions and processes them in parallel.

import dask.dataframe as dd

def process_video_with_dask(video_csv_dir, output_hdf_path):
    # Read all CSVs in the directory as a Dask DataFrame
    ddf = dd.read_csv(
        f"{video_csv_dir}/*.csv",
        dtype={"temperature_value": "float32"}
    )
    # Write to HDF (Dask handles partitioning automatically)
    ddf.to_hdf(output_hdf_path, key="temperature_data", mode="w")

Dask mimics Pandas' API, so you won't have to rewrite all your existing logic.

5. Split HDF Files into Smaller, Manageable Units

Storing an entire 10,000+ frame video in a single HDF file can still strain memory. Consider splitting each video's data into smaller chunks (e.g., 1000 frames per HDF key) or even separate HDF files for every 5000 frames.

For example:

chunk_counter = 0
for chunk in pd.read_csv(csv_file, chunksize=1000):
    chunk.to_hdf(
        output_hdf_path,
        key=f"frames_{chunk_counter}_to_{chunk_counter+999}",
        mode="a"
    )
    chunk_counter += 1000

Combining these strategies (especially incremental writing + data type optimization) should resolve your MemoryError issues. Start with the simplest fixes first, then move to Dask if you need to scale even further.

内容的提问来源于stack exchange,提问作者cedric0001

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:40:56