加密货币数据集多元回归拟合报错:无法将字符串转为浮点数
问题根源
你仅对目标变量Price做了逗号移除和类型转换,但特征列(Open/High/Low等)中同样存在带逗号的字符串格式数据,模型拟合时会尝试将特征矩阵X中的字符串转为浮点数,从而触发报错。
解决方案
1. 批量处理带逗号的数值列
定义通用处理函数,移除字符串中的逗号并转为浮点数,批量应用到需要的列:
def clean_numeric(x): if isinstance(x, str): x = x.replace(',', '') try: return float(x) except ValueError: return x return x # 处理Open/High/Low/Price列 cols_to_clean = ['Open', 'High', 'Low', 'Price'] df[cols_to_clean] = df[cols_to_clean].applymap(clean_numeric)
2. 处理特殊格式的特征列
- Vol.列:如果包含K(千)、M(百万)等后缀,需转换为实际数值:
def clean_volume(x): if isinstance(x, str): x = x.replace(',', '') if 'K' in x: return float(x.replace('K', '')) * 1000 elif 'M' in x: return float(x.replace('M', '')) * 1_000_000 elif 'B' in x: return float(x.replace('B', '')) * 1_000_000_000 else: try: return float(x) except: return None return x df['Vol.'] = df['Vol.'].apply(clean_volume)
- Change %列:移除%符号并转为浮点数:
df['Change %'] = df['Change %'].apply(lambda x: float(x.replace('%', '')) if isinstance(x, str) else x)
3. 验证数据类型并拟合模型
处理完成后,检查所有特征和目标列是否为数值型:
print(df[['Open', 'High', 'Low', 'Vol.', 'Change %', 'Price']].dtypes)
确认所有列均为float64或int64类型后,再执行模型拟合:
from sklearn.linear_model import LinearRegression X = df[['Open', 'High', 'Low', 'Vol.', 'Change %']] y = df['Price'] model = LinearRegression() model.fit(X, y)
内容的提问来源于stack exchange,提问作者Sweta
相关产品推荐
相关产品推荐

