LinearRegression拟合报错ValueError:字符串无法转浮点型求解决
Hey there! The error you're hitting happens because scikit-learn's Linear Regression model requires all input features to be numeric values, but your train_x includes the Section column which contains string identifiers like 06GA19960613—these can't be converted to floats automatically. Let's walk through the best ways to fix this:
Option 1: Drop the Section column (most straightforward if it's just an ID)
If the Section column is only a unique sample identifier and doesn't carry any predictive value for your target variable (pHSWS25), simply remove it from your feature set. Here's how to adjust your code:
import pandas as pd from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error train_data = pd.read_csv('train.csv') test_data = pd.read_csv('test.csv') print(train_data.head()) print('\nShape of training data :',train_data.shape) print('\nShape of testing data :',test_data.shape) # Drop both the target column AND the Section column from training features train_x = train_data.drop(columns=['pHSWS25', 'Section'], axis=1) train_y = train_data['pHSWS25'] # Don't forget to drop Section from test data too! test_x = test_data.drop(columns=['Section'], axis=1) print(train_x.head()) print(train_y.head()) # Now fit the model successfully model = LinearRegression() model.fit(train_x, train_y)
Option 2: Encode the Section column (if it has categorical meaning)
If Section represents a categorical feature (e.g., different geographic regions or sample groups that might impact the target), you'll need to convert these strings to numeric values. Two common approaches are:
Sub-option 2a: Label Encoding (for ordinal categories)
Use this if the Section values have an inherent order (e.g., "Group A" < "Group B" < "Group C"):
import pandas as pd from sklearn.linear_model import LinearRegression from sklearn.preprocessing import LabelEncoder train_data = pd.read_csv('train.csv') test_data = pd.read_csv('test.csv') # Initialize label encoder le = LabelEncoder() # Fit on training data and transform both train and test sets (critical to avoid data leakage!) train_data['Section'] = le.fit_transform(train_data['Section']) test_data['Section'] = le.transform(test_data['Section']) # Prepare features and target train_x = train_data.drop(columns=['pHSWS25'], axis=1) train_y = train_data['pHSWS25'] test_x = test_data print(train_x.head()) print(train_y.head()) model = LinearRegression() model.fit(train_x, train_y)
Sub-option 2b: One-Hot Encoding (for nominal categories)
Use this if Section values are unordered categories (e.g., different sample locations with no hierarchy):
import pandas as pd from sklearn.linear_model import LinearRegression from sklearn.preprocessing import OneHotEncoder from sklearn.compose import ColumnTransformer train_data = pd.read_csv('train.csv') test_data = pd.read_csv('test.csv') # Define column transformer to one-hot encode Section, keep other columns as-is ct = ColumnTransformer( transformers=[('section_encoder', OneHotEncoder(sparse=False, drop='first'), ['Section'])], remainder='passthrough' ) # Transform training features (fit and transform) train_x = ct.fit_transform(train_data.drop(columns=['pHSWS25'])) # Transform test features (only transform, using training fit) test_x = ct.transform(test_data) train_y = train_data['pHSWS25'] model = LinearRegression() model.fit(train_x, train_y)
Quick Tips:
- Always check your feature data types first with
train_x.dtypes—any column markedobjectis likely a string and needs handling. - When encoding, never fit the encoder on test data—always use the fit from the training set to avoid data leakage.
内容的提问来源于stack exchange,提问作者Siham MB

