使用MinMaxScaler处理DataFrame时遇ValueError问题求助
问题:MinMaxScaler归一化时序数据时触发ValueError
在对小时级时序数据预处理以应用K-means聚类时,使用MinMaxScaler做归一化出现如下错误:
Traceback (most recent call last): File ".venv\lib\site-packages\pandas\core\series.py", line 191, in wrapper raise TypeError(f"cannot convert the series to {converter}") TypeError: cannot convert the series to <class 'float'> The above exception was the direct cause of the following exception: Traceback (most recent call last): File ".venv\timesequence.py", line 210, in <module> matrix = pd.DataFrame(scaler.fit_transform(x_calls), columns=df_hours.columns, index=df_hours.index) File ".venv\lib\site-packages\sklearn\base.py", line 867, in fit_transform return self.fit(X, **fit_params).transform(X) File ".venv\lib\site-packages\sklearn\preprocessing\_data.py", line 420, in fit return self.partial_fit(X, y) File ".venv\lib\site-packages\sklearn\preprocessing\_data.py", line 457, in partial_fit X = self._validate_data( File ".venv\lib\site-packages\sklearn\base.py", line 577, in _validate_data X = check_array(X, input_name="X", **check_params) File ".venv\lib\site-packages\sklearn\utils\validation.py", line 856, in check_array array = np.asarray(array, order=order, dtype=dtype) File ".venv\lib\site-packages\pandas\core\generic.py", line 2064, in __array__ return np.asarray(self._values, dtype=dtype) ValueError: setting an array element with a sequence.
预处理及归一化代码如下:
#--------------------Preprocessing ds counter_ = 0 zero = 0 df_hours = pd.DataFrame({ 'Hour': [], 'SumView':[], 'CountStudent':[] }, dtype=object) while counter_ < 24: if (counter_ in sub_data_hour['Hour']): row = sub_data_hour.loc[(pd.to_numeric(sub_data_hour['Hour'], errors='coerce')) == counter_] df_hours.loc[len(df_hours.index)] = [counter_, row['SumView'], row['CountStudent']] else: df_hours.loc[len(df_hours.index)] = [counter_, zero, zero] counter_ += 1 #----------Normalize dataset------------ x_calls = df_hours.columns[2:] scaler = MinMaxScaler() matrix = pd.DataFrame(scaler.fit_transform(df_hours[x_calls]), columns=x_calls, index=df_hours.index)
错误原因
预处理循环中,当匹配到对应小时的行时,row['SumView']和row['CountStudent']返回的是Pandas Series对象(而非单个数值),导致df_hours的这两列混合了Series和标量值。MinMaxScaler要求输入为纯数值型二维数组,无法处理这种嵌套结构,因此抛出ValueError。
修复方案
1. 修正预处理逻辑,确保存入单个数值
将循环中赋值语句里的row['SumView']和row['CountStudent']改为取出Series的第一个元素(因为row是筛选后的单行DataFrame):
2. 移除不必要的object类型指定
初始化DataFrame时不要强制dtype=object,让Pandas自动推断数值类型,避免后续类型转换问题。
修复后的完整代码
#--------------------Preprocessing ds counter_ = 0 zero = 0 # 不指定dtype=object,让Pandas自动推断数值类型 df_hours = pd.DataFrame({ 'Hour': [], 'SumView': [], 'CountStudent': [] }) while counter_ < 24: if counter_ in sub_data_hour['Hour']: # 筛选对应小时的行 row = sub_data_hour.loc[pd.to_numeric(sub_data_hour['Hour'], errors='coerce') == counter_] # 用iloc[0]取出单个数值,而非Series对象 df_hours.loc[len(df_hours.index)] = [counter_, row['SumView'].iloc[0], row['CountStudent'].iloc[0]] else: df_hours.loc[len(df_hours.index)] = [counter_, zero, zero] counter_ += 1 #----------Normalize dataset------------ x_calls = df_hours.columns[2:] scaler = MinMaxScaler() # 此时df_hours[x_calls]是纯数值型DataFrame,可直接传入fit_transform matrix = pd.DataFrame(scaler.fit_transform(df_hours[x_calls]), columns=x_calls, index=df_hours.index)
额外优化建议
可以用更高效的Pandas方法替代while循环,比如reindex来补全24小时的数据,避免手动循环:
# 将sub_data_hour的Hour列转为数值型 sub_data_hour['Hour'] = pd.to_numeric(sub_data_hour['Hour'], errors='coerce') # 按Hour分组聚合 hourly_data = sub_data_hour.groupby('Hour')[['SumView', 'CountStudent']].sum() # 重新索引补全0-23小时,缺失值填充0 df_hours = hourly_data.reindex(range(24), fill_value=0).reset_index().rename(columns={'index': 'Hour'})
这种方法代码更简洁,运行效率更高,也能避免手动循环带来的错误。
内容的提问来源于stack exchange,提问作者teresaroserain
相关产品推荐
相关产品推荐

