机器学习入门遇字符串转浮点数错误:could not convert string to float: ' 8,400,000,000
解决"could not convert string to float: ' 8,400,000,000'"错误
问题根源
你的Price列中存在带千位分隔符逗号和前置空格的字符串值(比如' 8,400,000,000'),pandas默认将这类值识别为字符串类型,而sklearn的回归模型无法处理字符串格式的标签数据,因此训练时触发类型转换错误。
修复方案
1. 清理并转换Price列为数值类型
使用字符串处理方法去除空格和逗号,再转为float/int类型。
2. 修正缺失值处理逻辑
原代码中连续两次调用fillna,第二次会覆盖第一次的结果,需要合并处理或选择合适的填充方式。
3. 检查特征列的数值类型
确保Area等特征列也是数值类型,避免后续训练出现类似错误。
修改后的完整代码
import numpy as np import matplotlib.pyplot as plt import pandas as pd from sklearn.model_selection import train_test_split from sklearn.tree import DecisionTreeRegressor from sklearn.metrics import mean_absolute_error from sklearn import linear_model # 读取数据 df = pd.read_csv("housePrice.csv") # 查看缺失值和数据基本信息 print(df.isna().sum()) print(df.head()) print(df.describe()) print(df.info()) # -------------------------- 关键修复部分 -------------------------- # 清理Price列:去除空格、逗号,转换为数值类型 df['Price'] = df['Price'].str.strip().str.replace(',', '').astype(float) # 处理缺失值:先向前填充再向后填充(合并两次操作) df = df.fillna(method="ffill").fillna(method="bfill") # 确保特征列是数值类型(如果Area等列有字符串格式,同样处理) df['Area'] = df['Area'].str.strip().str.replace(',', '').astype(float) # Room、Parking、Warehouse如果是布尔或整数,确保类型正确 df['Room'] = pd.to_numeric(df['Room'], errors='coerce') df['Parking'] = df['Parking'].astype(int) df['Warehouse'] = df['Warehouse'].astype(int) # ------------------------------------------------------------------ # 提取特征和标签 x = df[["Area","Room","Parking","Warehouse"]] y = df['Price'] print(x.shape) print(y.shape) # 划分训练集和测试集 train_x, test_x, train_y, test_y = train_test_split(x , y , random_state=0, test_size=0.3) # 训练模型并评估 dt = DecisionTreeRegressor() dt.fit(train_x , train_y) pred_y = dt.predict(test_x) print("MAE:" , mean_absolute_error(test_y , pred_y))
额外说明
- 如果
Area列也存在带逗号的字符串格式,必须像处理Price一样清理转换,否则训练时同样会触发类型错误。 Parking和Warehouse如果是布尔值(比如True/False),可以直接转为int类型(1/0),模型能正常处理。- 若清理后仍有无法转换的值,
pd.to_numeric的errors='coerce'参数会将其转为NaN,后续的缺失值填充会处理这些情况。
内容的提问来源于stack exchange,提问作者Matin2002
相关产品推荐
相关产品推荐

