如何解决numpy.ndarray对象无append属性错误?附代码求助
Hey there! Let's break down why you're hitting this error and how to fix it. First off, that error almost always happens when someone tries to use the list-style .append() method on a numpy array (which doesn't support it). From looking at your code, the issue is likely in your custom HDF5DatasetWriter class, not the main feature extraction flow. Let's walk through the fixes step by step.
First: Spot Potential Issues in Your Main Code
Before diving into the core fix, let's tweak a couple of things in your main script that could cause unexpected behavior:
- Uncomment the ImageNet preprocessing: You've got
#image = imagenet_utils.preprocess_input(image)commented out. VGG16 is trained on ImageNet's normalized data, so skipping this will make your features less accurate. Uncomment that line. - Optimize batch image construction: Your current approach uses a list to append images then stacks them. While this works, we can make it more efficient and avoid type confusion by pre-allocating a numpy array upfront (we'll show this in the optimized code below).
Core Fix: Fix the HDF5DatasetWriter Class
The root of your error is almost certainly in how the add method is implemented in your HDF5DatasetWriter class. If it's trying to call .append() on a numpy array (instead of using buffer lists or proper numpy array writes), that's exactly what triggers this error.
Replace your existing hdf5datasetwriter.py code with this correct implementation:
import h5py import numpy as np class HDF5DatasetWriter: def __init__(self, dims, outputPath, dataKey="features", bufSize=1000): # Create the HDF5 file and datasets for features and labels self.db = h5py.File(outputPath, "w") self.data = self.db.create_dataset(dataKey, dims, dtype="float") self.labels = self.db.create_dataset("labels", (dims[0],), dtype="int") # Set up buffer to optimize writes self.bufSize = bufSize self.buffer = {"data": [], "labels": []} self.idx = 0 def add(self, rows, labels): # Add new data/labels to the in-memory buffer self.buffer["data"].extend(rows) self.buffer["labels"].extend(labels) # Flush buffer to disk if it's full if len(self.buffer["data"]) >= self.bufSize: self.flush() def flush(self): # Write buffer contents to HDF5 file current_batch_size = len(self.buffer["data"]) self.data[self.idx:self.idx+current_batch_size] = self.buffer["data"] self.labels[self.idx:self.idx+current_batch_size] = self.buffer["labels"] self.idx += current_batch_size # Reset buffer self.buffer = {"data": [], "labels": []} def storeClassLabels(self, classLabels): # Store human-readable class labels dt = h5py.special_dtype(vlen=str) label_dataset = self.db.create_dataset("label_names", (len(classLabels),), dtype=dt) label_dataset[:] = classLabels def close(self): # Flush any remaining data and close the file if len(self.buffer["data"]) > 0: self.flush() self.db.close()
Optimized Main Code (Optional but Recommended)
Here's your main script with the preprocessing uncommented and a more efficient batch image setup:
from keras.applications import VGG16 from keras.applications import imagenet_utils from keras.preprocessing.image import img_to_array from keras.preprocessing.image import load_img from sklearn.preprocessing import LabelEncoder from hdf5datasetwriter import HDF5DatasetWriter from imutils import paths import progressbar import argparse import random import numpy as np import os # Parse command line arguments ap = argparse.ArgumentParser() ap.add_argument("-d", "--dataset", required=True, help="Path to input dataset") ap.add_argument("-o", "--output", required=True, help="Path to output HDF5 file") ap.add_argument("-b","--batch_size", type=int, default=32, help="Batch size of images to process") ap.add_argument("-s","--buffer_size", type=int, default=1000, help="Feature extraction buffer size") args = vars(ap.parse_args()) bs = args["batch_size"] # Load and shuffle image paths print("[INFO] loading images...") imagePaths = list(paths.list_images(args["dataset"])) random.shuffle(imagePaths) # Extract and encode labels labels = [p.split(os.path.sep)[-2] for p in imagePaths] le = LabelEncoder() labels = le.fit_transform(labels) # Load pre-trained VGG16 (without top classification layer) print("[INFO] loading network...") model = VGG16(weights="imagenet", include_top=False) # Initialize HDF5 writer dataset = HDF5DatasetWriter((len(imagePaths), 512*7*7), args["output"], dataKey="features", bufSize=args["buffer_size"]) dataset.storeClassLabels(le.classes_) # Set up progress bar widgets = ["Extracting features:", progressbar.Percentage(), " ", progressbar.Bar(), " ", progressbar.ETA()] pbar = progressbar.ProgressBar(maxval=len(imagePaths), widgets=widgets).start() # Process images in batches for i in np.arange(0, len(imagePaths), bs): batchPaths = imagePaths[i:i + bs] batchLabels = labels[i:i+bs] # Pre-allocate numpy array for batch images (more efficient) batchImages = np.zeros((len(batchPaths), 224, 224, 3), dtype="float32") for (j, imagePath) in enumerate(batchPaths): # Load and preprocess image image = load_img(imagePath, target_size=(224,224)) image = img_to_array(image) image = np.expand_dims(image, axis=0) image = imagenet_utils.preprocess_input(image) # Add to batch array batchImages[j] = image # Extract features features = model.predict(batchImages, batch_size=bs) # Flatten features features = features.reshape((features.shape[0], 512*7*7)) # Add to HDF5 dataset dataset.add(features, batchLabels) pbar.update(i) # Cleanup dataset.close() pbar.finish()
Why Your Previous Fixes Didn't Work
You tried using np.append() and np.concatenate() in your main code, but the error wasn't coming from there—it was hidden in the HDF5DatasetWriter class's add method. The original implementation was probably trying to call .append() on a numpy array (a list method that numpy doesn't have). The fixed class uses an in-memory list buffer instead, then writes batches to the HDF5 file properly.
内容的提问来源于stack exchange,提问作者Mohamed Witti Adou

