如何高效保存图片数组及关联信息?以汽车爬取场景为例
Hey there! Great job getting the car images sorted out already—adding the parameter info in a clean, accessible structure like you want is totally achievable. Let’s break down how to build that dataset structure you described, where you can easily extract info with dataset['info'] or unpack with x, y = dataset.
Step 1: Define a Clean Single-Sample Structure
First, for individual car entries (one image + its parameters), using a namedtuple or dataclass works perfectly—they support both attribute/dict-style access and unpacking. Here’s how to use namedtuple:
from collections import namedtuple # Define the structure for a single car sample CarSample = namedtuple('CarSample', ['image', 'info']) # Example usage when you scrape a car # Assume `image_array` is your image data (numpy array, PIL Image, etc.) # `car_params` is your list of parameters like make, horsepower, etc. image_array = [255, 203, 145, ...] # Replace with actual image data car_params = ['Audi', '355 HP', 'Sedan', '2024'] # Create a sample instance sample = CarSample(image=image_array, info=car_params) # Access data in multiple ways print(sample['image']) # Dict-style access print(sample.info) # Attribute-style access # Unpack directly img, params = sample print(img) print(params)
If you prefer mutable structures (in case you need to edit data later), use a dataclass instead:
from dataclasses import dataclass @dataclass class CarSample: image: list # Or numpy.ndarray/PIL.Image depending on your data info: list # Usage is identical to the namedtuple example above
Step 2: Build a Dataset Container for Bulk Data
For a full dataset with multiple samples, create a custom class that mimics the behavior you want—dict-style access to all images/infos, unpacking support, and a clean print output:
class CarDataset: def __init__(self): self._images = [] self._infos = [] def add_sample(self, image, info): """Add a single car sample to the dataset""" self._images.append(image) self._infos.append(info) def __getitem__(self, key): """Enable dict-style access like dataset['image']""" if key == 'image': return self._images elif key == 'info': return self._infos else: raise KeyError(f"Unsupported key: {key}") def __iter__(self): """Enable unpacking like all_images, all_infos = dataset""" yield self._images yield self._infos def __repr__(self): """Custom print output to match your desired format""" # Truncate long lists for readability truncated_imgs = self._images[:3] + ['...'] if len(self._images) > 3 else self._images truncated_infos = self._infos[:3] + ['...'] if len(self._infos) > 3 else self._infos return f"{{ 'image': ({truncated_imgs}), 'info': ({truncated_infos}) }}"
Using the Dataset:
# Initialize the dataset dataset = CarDataset() # Add samples as you scrape them dataset.add_sample([255, 203, 145, ...], ['Audi', '355 HP', ...]) dataset.add_sample([120, 150, 180, ...], ['BMW', '320 HP', ...]) dataset.add_sample([90, 100, 110, ...], ['Mercedes', '380 HP', ...]) # Access all images/infos print(dataset['image']) print(dataset['info']) # Unpack the entire dataset all_images, all_infos = dataset # Print the dataset (matches your desired output) print(dataset) # Output: { 'image': ([255, 203, 145, ...], [120, 150, 180, ...], [90, 100, 110, ...], ...), 'info': (['Audi', '355 HP', ...], ['BMW', '320 HP', ...], ['Mercedes', '380 HP', ...], ...) }
Step 3: Efficiently Save & Load the Dataset
For long-term storage, use serialization that’s efficient for your data size:
Small to Medium Datasets: Pickle
Pickle is easy and works with our custom classes out of the box:
import pickle # Save the dataset with open('car_dataset.pkl', 'wb') as f: pickle.dump(dataset, f) # Load the dataset later with open('car_dataset.pkl', 'rb') as f: loaded_dataset = pickle.load(f) # Use it just like before print(loaded_dataset['info'])
Large Datasets: HDF5
If you’re dealing with thousands of images, HDF5 is more memory-efficient (use h5py):
import h5py # Save (note: for string data in 'info', we need to store as variable-length strings) with h5py.File('car_dataset.h5', 'w') as f: # Store images (assuming they're numpy arrays) f.create_dataset('images', data=dataset['image']) # Store info as variable-length strings dt = h5py.special_dtype(vlen=str) f.create_dataset('info', data=dataset['info'], dtype=dt) # Load with h5py.File('car_dataset.h5', 'r') as f: loaded_images = f['images'][:] loaded_infos = f['info'][:]
Bonus: Make It Iterable for Batch Processing
If you want to loop through individual samples (like PyTorch/TensorFlow datasets), add these methods to CarDataset:
def __len__(self): """Return the number of samples""" return len(self._images) def __getitem__(self, idx): """Override to support both key-based and index-based access""" if isinstance(idx, str): return super().__getitem__(idx) elif isinstance(idx, int): return CarSample(self._images[idx], self._infos[idx])
Now you can do for sample in dataset: or access dataset[0] to get a single CarSample instance.
内容的提问来源于stack exchange,提问作者Nicolas Gervais

