对Numpy ndarray执行归一化操作触发TypeError错误的技术问询
Hey there, let's break down why you're running into this error and fix it step by step!
The Root Cause
When you read your CSV data using csv.reader and convert it to a NumPy array, all the values are stored as string types (NumPy calls this a "flexible type" like dtype='<U10'). Numerical operations like calculating means, minima, or peak-to-peak values (ptp()) require numeric data types (int/float), so trying to run these operations on string arrays triggers that TypeError.
Solution 1: Convert to Numeric Type When Creating the Array
The simplest fix is to tell NumPy to parse the data as numeric values right when you create the array. Replace your line:
data = np.array(x)
with:
data = np.array(x, dtype=np.float64)
This forces all elements to be 64-bit floats, which supports all the math operations you need.
Solution 2: Convert the Feature Array Later
If you need to keep the original array (e.g., the first column is a classification label you want as a string), you can just convert the feature subset to numeric types:
X_raw = data[:,1:13].astype(np.float64)
Full Corrected Code (With Both Normalization Options)
Here's your code updated to work, including the standard (X-μ)/σ normalization you originally wanted:
import csv import numpy as np import urllib.request # Download the dataset url = 'http://archive.ics.uci.edu/ml/machine-learning-databases/wine/wine.data' urllib.request.urlretrieve(url,'F:/Python/Wine Dataset/wine_data') # Read and parse the data filename = 'F:/Python/Wine Dataset/wine_data' raw_data = open(filename,'rt') reader = csv.reader(raw_data) x = list(reader) # Convert to numeric array immediately data = np.array(x, dtype=np.float64) # Split labels and features y = data[:,0] X_raw = data[:,1:13] # Option 1: Min-max normalization (your original simplified attempt) X_minmax = (X_raw - X_raw.min(0)) / X_raw.ptp(0) print("Min-max normalized data:\n", X_minmax[:5]) # Print first 5 rows to check # Option 2: Standardization (X-μ)/σ (your original desired method) mu = X_raw.mean(0) sigma = X_raw.std(0) X_standardized = (X_raw - mu) / sigma print("\nStandardized data:\n", X_standardized[:5])
Bonus: Use NumPy's Built-in CSV Reader
For even cleaner code, you can skip csv.reader entirely and use np.genfromtxt, which automatically handles type conversion:
data = np.genfromtxt(filename, delimiter=',')
This reads the CSV directly into a numeric NumPy array, no extra steps needed!
内容的提问来源于stack exchange,提问作者W. Roberts

