You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LinearRegression拟合报错ValueError:字符串无法转浮点型求解决

Fixing ValueError: invalid literal for float() in Linear Regression

Hey there! The error you're hitting happens because scikit-learn's Linear Regression model requires all input features to be numeric values, but your train_x includes the Section column which contains string identifiers like 06GA19960613—these can't be converted to floats automatically. Let's walk through the best ways to fix this:

Option 1: Drop the Section column (most straightforward if it's just an ID)

If the Section column is only a unique sample identifier and doesn't carry any predictive value for your target variable (pHSWS25), simply remove it from your feature set. Here's how to adjust your code:

import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error

train_data = pd.read_csv('train.csv')
test_data = pd.read_csv('test.csv')

print(train_data.head())
print('\nShape of training data :',train_data.shape)
print('\nShape of testing data :',test_data.shape)

# Drop both the target column AND the Section column from training features
train_x = train_data.drop(columns=['pHSWS25', 'Section'], axis=1)
train_y = train_data['pHSWS25']

# Don't forget to drop Section from test data too!
test_x = test_data.drop(columns=['Section'], axis=1)

print(train_x.head())
print(train_y.head())

# Now fit the model successfully
model = LinearRegression()
model.fit(train_x, train_y)

Option 2: Encode the Section column (if it has categorical meaning)

If Section represents a categorical feature (e.g., different geographic regions or sample groups that might impact the target), you'll need to convert these strings to numeric values. Two common approaches are:

Sub-option 2a: Label Encoding (for ordinal categories)

Use this if the Section values have an inherent order (e.g., "Group A" < "Group B" < "Group C"):

import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import LabelEncoder

train_data = pd.read_csv('train.csv')
test_data = pd.read_csv('test.csv')

# Initialize label encoder
le = LabelEncoder()

# Fit on training data and transform both train and test sets (critical to avoid data leakage!)
train_data['Section'] = le.fit_transform(train_data['Section'])
test_data['Section'] = le.transform(test_data['Section'])

# Prepare features and target
train_x = train_data.drop(columns=['pHSWS25'], axis=1)
train_y = train_data['pHSWS25']
test_x = test_data

print(train_x.head())
print(train_y.head())

model = LinearRegression()
model.fit(train_x, train_y)

Sub-option 2b: One-Hot Encoding (for nominal categories)

Use this if Section values are unordered categories (e.g., different sample locations with no hierarchy):

import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import OneHotEncoder
from sklearn.compose import ColumnTransformer

train_data = pd.read_csv('train.csv')
test_data = pd.read_csv('test.csv')

# Define column transformer to one-hot encode Section, keep other columns as-is
ct = ColumnTransformer(
    transformers=[('section_encoder', OneHotEncoder(sparse=False, drop='first'), ['Section'])],
    remainder='passthrough'
)

# Transform training features (fit and transform)
train_x = ct.fit_transform(train_data.drop(columns=['pHSWS25']))
# Transform test features (only transform, using training fit)
test_x = ct.transform(test_data)

train_y = train_data['pHSWS25']

model = LinearRegression()
model.fit(train_x, train_y)

Quick Tips:

  • Always check your feature data types first with train_x.dtypes—any column marked object is likely a string and needs handling.
  • When encoding, never fit the encoder on test data—always use the fit from the training set to avoid data leakage.

内容的提问来源于stack exchange,提问作者Siham MB

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:43:18