使用回归模型预测时遇ValueError:无法将'veh2312'转为整数
回归模型预测时的ValueError问题解决
问题描述
使用随机森林回归模型做预测时触发错误:ValueError: invalid literal for int() with base 10: 'veh2312',多次尝试修复无效,寻求解决方法。
错误栈
---> 47 sample_input = df[(df['vehicle_id'] == label_encoder.transform([vehicle_id])[0]) & (df['time'] == time_step)] /usr/local/lib/python3.10/dist-packages/sklearn/utils/_array_api.py in astype(self, x, dtype, copy, casting) 66 def astype(self, x, dtype, *, copy=True, casting="unsafe"): 67 # astype is not defined in the top level NumPy namespace ---> 68 return x.astype(dtype, copy=copy, casting=casting ValueError: invalid literal for int() with base 10: 'veh2312'
完整代码
#'time' column to Unix timestamp df['time'] = pd.to_datetime(df['time']).astype(int) // 10**9 label_encoder = LabelEncoder() df['vehicle_id'] = label_encoder.fit_transform(df['vehicle_id']) X = df[['time', 'vehicle_id', 'vehicle_lat', 'vehicle_lon', 'cell_lat', 'cell_lon']] # Features y = df.groupby(['time', 'vehicle_id'])['cell_id'].nunique().reset_index(name='cell_id_count') df_merged = pd.merge(df, y, on=['time', 'vehicle_id']) y = df_merged['cell_id_count'] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) rf_regressor = RandomForestRegressor(n_estimators=10, random_state=42) rf_regressor.fit(X_train, y_train) y_pred = rf_regressor.predict(X_test) #input vehicle_id = 'veh2312' time_step = 2194 #filter DataFrame based on 'vehicle_id' and 'time' sample_input = df[(df['vehicle_id'] == label_encoder.transform([vehicle_id])[0]) & (df['time'] == time_step)] if len(sample_input) > 0: sample_features = sample_input[['time', 'vehicle_id', 'vehicle_lat', 'vehicle_lon', 'cell_lat', 'cell_lon']] sample_predictions = rf_regressor.predict(sample_features) print(f"Predicted number of cell_ids for vehicle '{vehicle_id}' at time step {time_step}: {sample_predictions[0]:.0f}") else: print(f"No data found for vehicle '{vehicle_id}' at time step {time_step}.")
样本数据集
cell_id cell_lat cell_lon time vehicle_lat vehicle_lon vehicle_id distance 2665 43.720118 10.422051 2194 43.70994621 10.42621903 veh2312 1205 2682 43.695238 10.433212 2194 43.70994621 10.42621903 veh2312 1787 20020711 43.716044 10.430057 2194 43.70994621 10.42621903 veh2312 791 27025350 43.710199 10.4247325 2194 43.70994621 10.42621903 veh2312 167 118817885 43.712409 10.424214 2194 43.70994621 10.42621903 veh2312 349 118825175 43.706199 10.431691 2194 43.70994621 10.42621903 veh2312 731 118840263 43.706233 10.432086 2194 43.70994621 10.42621903 veh2312 766
解决方案
错误原因
- 时间列处理逻辑错误:代码错误地将代表时间步长的
time列转换为Unix时间戳,导致df['time']的值与输入的time_step=2194完全不匹配;若数据读取时存在列错位,会导致time列混入vehicle_id的字符串值。 - 数据类型不匹配:当
df['time']包含字符串(如'veh2312')时,执行df['time'] == time_step(整数)时,Pandas会尝试将字符串转换为整数,触发ValueError。
修复步骤
- 修正时间列处理:如果
time列是时间步长而非日期时间,删除错误的Unix时间戳转换代码,保留原始数值:# 移除该行错误代码 # df['time'] = pd.to_datetime(df['time']).astype(int) // 10**9 - 确保数据读取正确:读取数据时指定正确的分隔符和表头,保证
time列为整数类型,vehicle_id列为字符串类型:df = pd.read_csv('your_data.csv', sep='\s+', header=0) # 强制转换time列为整数 df['time'] = df['time'].astype(int) - 验证筛选逻辑:筛选样本前确认数据类型一致,避免隐式转换错误:
encoded_vehicle_id = label_encoder.transform([vehicle_id])[0] # 确保time_step与df['time']类型匹配 sample_input = df[(df['vehicle_id'] == encoded_vehicle_id) & (df['time'] == int(time_step))]
内容的提问来源于stack exchange,提问作者user1566490
相关产品推荐
相关产品推荐

