如何在Sklearn随机森林预测中保留测试数据索引并导出至Excel?
解决方法
1. 关联预测结果与测试数据的索引、坐标
核心是把预测值和测试集的索引、坐标、真实值整合到同一个结构化数据里,用Pandas的DataFrame最方便——不管你的测试集是NumPy数组还是DataFrame,都能轻松处理。
代码示例:
import pandas as pd # 情况1:X_test本身是带索引和坐标列的DataFrame result_df = pd.DataFrame({ '原始索引': X_test.index, # 直接取测试集的原始索引 '经度': X_test['lon'], # 替换成你实际的经度列名 '纬度': X_test['lat'], # 替换成你实际的纬度列名 '真实值': y_test.values if isinstance(y_test, pd.Series) else y_test, '预测值': y_pred }) # 情况2:X_test是NumPy数组(比如用train_test_split直接拆分的) # 先从原始数据里拿测试集的索引、坐标(假设拆分前原始数据是df) # test_indices = df.iloc[X_test.index].index # 或者你拆分时手动记录的测试集索引 # result_df = pd.DataFrame({ # '原始索引': test_indices, # '经度': df.iloc[X_test.index]['lon'], # '纬度': df.iloc[X_test.index]['lat'], # '真实值': y_test, # '预测值': y_pred # })
2. 导出到Excel文件
用Pandas的to_excel直接导出,比np.savetxt更适合带结构化信息的结果:
# 导出时可以关闭默认的行索引,避免重复 result_df.to_excel('预测结果带坐标.xlsx', index=False)
3. 生成地图的实用思路
有了带坐标的结果,用这两个工具快速生成地图:
- Plotly Express:做交互式散点地图,一键出图:
import plotly.express as px fig = px.scatter_mapbox(result_df, lat='纬度', lon='经度', color='预测值', zoom=10, mapbox_style='carto-positron') fig.show() - Folium:做带自定义标记的地图,适合查看单条数据的细节:
import folium # 以测试集坐标的平均值为地图中心 m = folium.Map(location=[result_df['纬度'].mean(), result_df['经度'].mean()], zoom_start=10) # 给每个坐标点加标记,弹窗显示索引、真实值和预测值 for _, row in result_df.iterrows(): folium.Marker( location=[row['纬度'], row['经度']], popup=f"索引: {row['原始索引']}<br>真实值: {row['真实值']}<br>预测值: {row['预测值']}", icon=folium.Icon(color='blue' if row['预测值'] < row['真实值'] else 'red') ).add_to(m) # 保存为HTML文件,直接打开就能看 m.save('预测结果地图.html')
关键提醒
如果用train_test_split拆分数据,默认会返回NumPy数组,丢失原始索引。建议拆分前把数据保留为Pandas DataFrame,这样拆分后的X_test和y_test会自动保留原始索引,省得手动找对应关系。
内容的提问来源于stack exchange,提问作者hanie kalantar
相关产品推荐
相关产品推荐

