You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据预处理时遇「无法将字符串转换为浮点数」错误求助(新手)

Fixing "Could Not Convert String to Float" Error in Data Preprocessing

Hey there! I totally get how frustrating this error can be when you're just starting out with data preprocessing. Let's break down what's happening and fix it step by step.

Why This Error Happens

The MinMaxScaler from scikit-learn only works with numerical data (integers or floats). When you hit this error, it means one or more columns in your training_set are stored as strings (object type) instead of numerical values. Common culprits in stock price datasets are:

  • Volume columns with commas (like "1,234,567" instead of 1234567)
  • Missing values labeled as strings (like "N/A" or empty cells)
  • Typos or random non-numeric entries in otherwise numerical columns

Step-by-Step Fix

1. First, Spot the Problem Columns

Before fixing anything, let's check what data types we're dealing with. Add these lines right after loading your dataset:

# Check data types of the columns you're using
print(dataset_train.iloc[:, 1:6].dtypes)

Look for columns marked as object—those are the string columns causing the issue.

2. Clean String Columns (e.g., Volume with Commas)

Let's handle the most common case first: columns with commas (like Volume). If your Volume column is an object type, replace commas and convert it to float:

# Fix Volume column (adjust the column name if yours is different)
dataset_train['Volume'] = dataset_train['Volume'].str.replace(',', '').astype(float)

3. Convert All Columns to Numeric

To catch any other non-numeric entries, use pd.to_numeric to convert all target columns. This will turn unconvertable values into NaN (missing values) instead of crashing your code:

training_set = dataset_train.iloc[:, 1:6].apply(pd.to_numeric, errors='coerce')

4. Handle Missing Values

Now we need to deal with those NaN values. You can either drop rows with missing values or fill them with a reasonable value (like the column mean):

# Option 1: Fill missing values with column mean (great for preserving data)
training_set = training_set.fillna(training_set.mean())

# Option 2: Drop rows with any missing values (use only if you have plenty of data)
# training_set = training_set.dropna()

5. Scale the Data (Now It Should Work!)

Now that all your data is numeric, the scaler will run without errors:

from sklearn.preprocessing import MinMaxScaler
sc = MinMaxScaler(feature_range=(0, 1))
training_set_scaled = sc.fit_transform(training_set)

Full Modified Code

Putting it all together:

import numpy as np
import matplotlib.pyplot as plt
import pandas as pd

# Import training set
dataset_train = pd.read_csv('Google_Stock_Price_Train.csv')

# Check data types to identify problem columns
print(dataset_train.iloc[:, 1:6].dtypes)

# Clean Volume column (adjust if your problematic column is different)
dataset_train['Volume'] = dataset_train['Volume'].str.replace(',', '').astype(float)

# Convert columns to numeric, handle unconvertable values
training_set = dataset_train.iloc[:, 1:6].apply(pd.to_numeric, errors='coerce')

# Fill missing values with column mean
training_set = training_set.fillna(training_set.mean())

# Scale the preprocessed data
from sklearn.preprocessing import MinMaxScaler
sc = MinMaxScaler(feature_range=(0, 1))
training_set_scaled = sc.fit_transform(training_set)

If you still run into issues after this, double-check your dataset for any weird entries (like random text in numerical columns)—sometimes a stray "error" string can sneak in!

内容的提问来源于stack exchange,提问作者Max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:53:14