You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

线性回归训练报错:字符串转浮点数失败的原因与解决

问题描述

我正在开展1995-2020年美国贫困状况追踪项目,需要绘制线性回归散点图。执行以下代码时:

# Create a regression object.
regression = LinearRegression()  
# This is the regression object, which will be fit onto the training set.

# Fit the regression object onto the training set.
regression.fit(X_train, y_train)

出现错误:

ValueError: could not convert string to float: 'New Hampshire'

我猜测报错原因是New Hampshire这类州名包含空格(多单词),请问该猜测是否正确?如何将这些州名中的空格替换为下划线?涉及的州名包括:New Hampshire、New Jersey、New Mexico、New York、North Carolina、North Dakota、Rhode Island、South Carolina、South Dakota、West Virginia。

完整报错堆栈信息如下:

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
Cell In[55], line 3
      1 # Fit the regression object onto the training set.
----> 3 regression.fit(X_train, y_train)

File ~/anaconda3/lib/python3.11/site-packages/sklearn/linear_model/_base.py:648, in LinearRegression.fit(self, X, y, sample_weight)
    644 n_jobs_ = self.n_jobs
    646 accept_sparse = False if self.positive else ["csr", "csc", "coo"]
--> 648 X, y = self._validate_data(
    649     X, y, accept_sparse=accept_sparse, y_numeric=True, multi_output=True
    650 )
    652 sample_weight = _check_sample_weight(
    653     sample_weight, X, dtype=X.dtype, only_non_negative=True
    654 )
    656 X, y, X_offset, y_offset, X_scale = _preprocess_data(
    657     X,
    658     y,
   (...)
    661     sample_weight=sample_weight,
    662 )

File ~/anaconda3/lib/python3.11/site-packages/sklearn/base.py:584, in BaseEstimator._validate_data(self, X, y, reset, validate_separately, **check_params)
    582         y = check_array(y, input_name="y", **check_y_params)
    583     else:
--> 584         X, y = check_X_y(X, y, **check_params)
    585     out = X, y
    587 if not no_val_X and check_params.get("ensure_2d", True):

File ~/anaconda3/lib/python3.11/site-packages/sklearn/utils/validation.py:1106, in check_X_y(X, y, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, multi_output, ensure_min_samples, ensure_min_features, y_numeric, estimator)
   1101         estimator_name = _check_estimator_name(estimator)
   1102     raise ValueError(
   1103         f"{estimator_name} requires y to be passed, but the target y is None"
   1104     )
--> 1106 X = check_array(
   1107     X,
   1108     accept_sparse=accept_sparse,
   1109     accept_large_sparse=accept_large_sparse,
   1110     dtype=dtype,
   1111     order=order,
   1112     copy=copy,
   1113     force_all_finite=force_all_finite,
   1114     ensure_2d=ensure_2d,
   1115     allow_nd=allow_nd,
   1116     ensure_min_samples=ensure_min_samples,
   1117     ensure_min_features=ensure_min_features,
   1118     estimator=estimator,
   1119     input_name="X",
   1120 )
   1122 y = _check_y(y, multi_output=multi_output, y_numeric=y_numeric, estimator=estimator)
   1124 check_consistent_length(X, y)

File ~/anaconda3/lib/python3.11/site-packages/sklearn/utils/validation.py:879, in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name)
    877         array = xp.astype(array, dtype, copy=False)
    878     else:
--> 879         array = _asarray_with_order(array, order=order, dtype=dtype, xp=xp)
    880 except ComplexWarning as complex_warning:
    881     raise ValueError(
    882         "Complex data not supported\n{}\n".format(array)
    883     ) from complex_warning

File ~/anaconda3/lib/python3.11/site-packages/sklearn/utils/_array_api.py:185, in _asarray_with_order(array, dtype, order, copy, xp)
    182     xp, _ = get_namespace(array)
    183 if xp.__name__ in {"numpy", "numpy.array_api"}:
    184     # Use NumPy API to support order
--> 185     array = numpy.asarray(array, order=order, dtype=dtype)
    186     return xp.asarray(array, copy=copy)
    187 else:

ValueError: could not convert string to float: 'New Hampshire'
解答

你的猜测不正确。报错的根本原因不是州名里的空格,而是Scikit-learn的线性回归模型无法直接处理字符串类型的特征——无论字符串有没有空格,只要是文本格式,模型都无法将其转换为数值进行计算。

一、如果确实需要替换空格为下划线(仅处理文本格式,不解决模型适配问题)

如果你只是想修改州名的格式,用Pandas处理的话可以这样做:
假设你的州名列是state,可以用str.replace()方法批量替换:

import pandas as pd

# 假设你的数据集是df
df['state'] = df['state'].str.replace(' ', '_')

这条代码会把所有州名里的空格替换成下划线,比如New Hampshire变成New_Hampshire,你列出的所有多单词州名都适用。

二、解决线性回归报错的正确方法

要让模型能处理州名这类分类特征,需要对其进行数值编码,常用的两种方式:

  • 独热编码(One-Hot Encoding)
    适合没有顺序关系的分类特征(州名之间没有等级差异),用pd.get_dummies()或者Scikit-learn的OneHotEncoder实现:
# 用Pandas快速实现独热编码
df_encoded = pd.get_dummies(df, columns=['state'], drop_first=True)
# drop_first=True避免多重共线性,可选

之后用编码后的数据集df_encoded来拆分X_train和y_train即可。

  • 标签编码(Label Encoding)
    不建议用于州名这类无顺序的分类特征,因为会给州名赋予无意义的数值顺序,可能干扰模型结果,仅适用于有明确顺序的分类变量。

总结

替换空格无法解决你的报错问题,必须对州名这类字符串特征进行数值编码,独热编码是最适合的方案。

内容的提问来源于stack exchange,提问作者EC Cotterman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 08:36:17