Python 3中使用to_hdf存储多视频DataFrame触发MemoryError
Hey there, let's tackle this memory issue you're facing with processing those temperature image CSVs into HDF files. I've dealt with similar large-scale data processing headaches before, so here are some practical, actionable solutions to try out:
Instead of loading an entire video's worth of CSVs into a single DataFrame before saving, process and append data in chunks. This keeps only a small portion of your data in memory at any time.
import pandas as pd from pathlib import Path def process_single_video(video_csv_dir, output_hdf_path): # Iterate through each CSV in the video's directory for csv_file in Path(video_csv_dir).glob("*.csv"): # Read CSV in chunks (adjust chunksize based on your memory capacity) for chunk in pd.read_csv(csv_file, chunksize=1000): # Append chunk to HDF file chunk.to_hdf( output_hdf_path, key="temperature_frames", mode="a", append=True, format="table" # Required for appending to HDF ) # Example usage: process each video one by one video_dirs = ["video_1_csvs", "video_2_csvs", ..., "video_10_csvs"] for idx, dir_path in enumerate(video_dirs): output_hdf = f"video_{idx+1}_data.h5" process_single_video(dir_path, output_hdf)
Python's garbage collector doesn't always free up memory immediately after variables go out of scope. Explicitly cleaning up after each video ensures you don't carry over unused data to the next processing step.
import gc # After processing each video: del chunk # Delete any large variables from the current video gc.collect() # Trigger manual garbage collection
Add this at the end of your video processing loop to wipe the slate clean for the next video.
Temperature data rarely needs the default float64 precision. Downcasting to smaller numeric types can cut your memory usage by 50% or more.
# Specify dtypes when reading CSVs to save memory upfront df = pd.read_csv( csv_file, dtype={ "temperature_value": "float32", # Use float32 instead of float64 "frame_id": "int32" # Use int32 instead of int64 if frame counts don't exceed 2 billion } ) # Check memory usage to verify savings print(df.info(memory_usage="deep"))
If your data is way too big for even chunked Pandas, Dask is a drop-in replacement that handles datasets larger than your available RAM. It splits data into manageable partitions and processes them in parallel.
import dask.dataframe as dd def process_video_with_dask(video_csv_dir, output_hdf_path): # Read all CSVs in the directory as a Dask DataFrame ddf = dd.read_csv( f"{video_csv_dir}/*.csv", dtype={"temperature_value": "float32"} ) # Write to HDF (Dask handles partitioning automatically) ddf.to_hdf(output_hdf_path, key="temperature_data", mode="w")
Dask mimics Pandas' API, so you won't have to rewrite all your existing logic.
Storing an entire 10,000+ frame video in a single HDF file can still strain memory. Consider splitting each video's data into smaller chunks (e.g., 1000 frames per HDF key) or even separate HDF files for every 5000 frames.
For example:
chunk_counter = 0 for chunk in pd.read_csv(csv_file, chunksize=1000): chunk.to_hdf( output_hdf_path, key=f"frames_{chunk_counter}_to_{chunk_counter+999}", mode="a" ) chunk_counter += 1000
Combining these strategies (especially incremental writing + data type optimization) should resolve your MemoryError issues. Start with the simplest fixes first, then move to Dask if you need to scale even further.
内容的提问来源于stack exchange,提问作者cedric0001

