Python中复杂数据结构存储选型及内存优化方案咨询
Hey there! Let's tackle your memory issue head-on—your current setup with nested lists of custom objects makes sense for appending, but those large collections can quickly eat up RAM. Here's a breakdown of actionable optimizations, including whether switching to arrays makes sense for your use case:
First: Cut Down Per-Object Overhead
The biggest low-hanging fruit here is reducing the memory footprint of your van, bus, and motorcycle instances. By default, every Python object has a __dict__ (a dynamic dictionary) to store attributes, which adds significant overhead—especially when you have thousands of instances.
Fix this by adding __slots__ to each class. This tells Python to use a fixed-size array instead of a dynamic dictionary for attributes, cutting memory usage by 30-50% for each object:
class van: __slots__ = ['attr1', 'attr2'] # List ALL attributes your class uses def __init__(self): self.attr1 = 0 self.attr2 = "example" class bus: __slots__ = ['capacity', 'fuel_type'] def __init__(self): self.capacity = 50 self.fuel_type = "diesel" # Same for motorcycle
This is trivial to implement and will give you an immediate memory win without changing your core data structure.
Should You Switch to Arrays? Short Answer: It Depends
You mentioned arrays being more efficient for read-only data—but let's clarify:
- Python's built-in
arraymodule: Only works with basic data types (ints, floats, bytes), not custom objects. So this won't help you storevan/businstances directly. - NumPy arrays: If you store custom objects in a NumPy array, it's essentially just a contiguous block of pointers to those objects—no better than a list in terms of memory overhead. The real benefit comes when you can extract object attributes into numerical arrays instead of storing full objects.
For example, if all your objects have numerical attributes, replace the object instances with structured NumPy arrays or pandas DataFrames:
import numpy as np # Define a dtype that matches your van's attributes van_dtype = [('speed', 'f4'), ('weight', 'f4'), ('color', 'U10')] total_vans = 7 * 5000 # 7 samples × 5000 vans each all_vans = np.empty(total_vans, dtype=van_dtype) # Fill the array directly (no van instances!) for sample_idx in range(7): start = sample_idx * 5000 end = start + 5000 all_vans[start:end]['speed'] = np.random.rand(5000) * 100 all_vans[start:end]['weight'] = np.random.rand(5000) * 2000 all_vans[start:end]['color'] = "blue"
This eliminates all Python object overhead—NumPy stores data in contiguous memory blocks, which is way more efficient. If your analysis only needs the object attributes (not the full instances), this is a game-changer.
Optimize Your Dataset Structure
Your current all_data = [ [vans, buses, mcycles], ... ] setup creates a lot of nested lists, each with their own internal overhead (pointer arrays, size tracking, etc.). Instead, group all objects of the same type into single global collections, and track sample boundaries with indices:
all_vans = [] all_buses = [] all_mcycles = [] sample_metadata = [] for i in range(7): # Generate objects for the sample van_batch = [van() for _ in range(5000)] bus_batch = [bus() for _ in range(2000)] mcycle_batch = [motorcycle() for _ in range(3000)] # Extend global collections (cheaper than creating new lists per sample) all_vans.extend(van_batch) all_buses.extend(bus_batch) all_mcycles.extend(mcycle_batch) # Track where this sample's objects start/end in each global list sample_metadata.append({ 'van': (len(all_vans) - 5000, len(all_vans)), 'bus': (len(all_buses) - 2000, len(all_buses)), 'mcycle': (len(all_mcycles) - 3000, len(all_mcycles)) }) # When analyzing sample 3 (index 2): start_van, end_van = sample_metadata[2]['van'] sample_vans = all_vans[start_van:end_van]
This reduces the number of list objects you're storing and cuts down on redundant overhead from nested lists.
For Extreme Memory Pressure: Disk-Based Storage
If even after the above optimizations your data is too big for RAM, use disk-based storage that lets you load only what you need for analysis:
- NumPy Memmap: Creates a memory-mapped array that lives on disk but acts like an in-memory array. You can slice it to load specific samples without reading the entire dataset.
- HDF5 (via h5py): A hierarchical format perfect for large datasets. You can store each class's data as a separate dataset, and store sample indices in another. When analyzing, you read only the slices you need from disk.
Quick Priority List for Action
- Add
__slots__to your classes: Easiest, fastest win. - Restructure to global collections + indices: Reduces list overhead without changing your object model.
- Switch to numerical arrays/DataFrames: If your analysis doesn't need full object instances, this gives maximum memory efficiency.
- Disk-based storage: For datasets that are too large to fit in RAM even after optimizations.
内容的提问来源于stack exchange,提问作者Csenger Kovácsházi

