代码报错‘could not convert string to float: 'age'’的原因排查求助
Hey there, let's break down what's causing this error and fix it step by step.
The Root Cause
That error pops up because when you use np.loadtxt("car.csv", delimiter=",") to load your data, it's trying to read the header row (the first line with column names like 'age') as numerical data. Since strings can't be converted to floats, Python throws that error. Also, your code has redundant work—you're loading the CSV once with Pandas and again with NumPy, which isn't necessary and adds confusion.
Solutions
Let's go over two ways to fix this, with the second being the more robust option for CSV data:
1. Skip the Header with NumPy
If you want to stick with np.loadtxt, just add the skiprows=1 parameter to skip the first header line:
dataset = np.loadtxt("car.csv", delimiter=",", skiprows=1)
This tells NumPy to ignore the first line and start reading from the actual numerical data rows.
2. Use Pandas for Cleaner Data Handling (Recommended)
Since you already imported Pandas, let's lean into it—it's designed for tabular data like CSV files and handles headers automatically. Here's your revised, streamlined code:
import matplotlib.pyplot as plt import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.model_selection import cross_val_score from sklearn.model_selection import KFold from sklearn.pipeline import Pipeline from sklearn.preprocessing import MinMaxScaler from tensorflow.python.keras.models import Sequential from tensorflow.python.keras.layers import Dense from tensorflow.python.keras.wrappers.scikit_learn import KerasRegressor import os # Load CSV with Pandas (automatically recognizes headers) df = pd.read_csv('car.csv') # Replace commas with dots—make sure to assign back to df! df = df.replace(",", ".", regex=True) # Extract features (first 5 columns) and target (6th column) x = df.iloc[:, 0:5].values y = df.iloc[:, 5].values.reshape(-1, 1) # Use separate scalers for features and target to avoid scaling conflicts x_scaler = MinMaxScaler() xscale = x_scaler.fit_transform(x) y_scaler = MinMaxScaler() yscale = y_scaler.fit_transform(y) print("Scaling completed successfully!")
A couple key fixes here:
df.replace(",", ".")needs to be assigned back todf—Pandas methods don't modify the original DataFrame by default.- Using separate
MinMaxScalerinstances forxandyensures their scaling rules don't interfere with each other, which will make inverse scaling (if you need to convert predictions back to original values) much easier.
Quick Recap
The main issue was not accounting for the CSV header when loading data. Pandas simplifies this process significantly, so it's the better choice for most tabular data tasks.
内容的提问来源于stack exchange,提问作者Kalina Scarbrough

