构建推荐系统时持续触发IndexError索引越界错误,求排查解决
解决IndexError问题的步骤
首先看你代码里的明显错误:
- 重置索引的那行变量名拼写错误:
sampld_data = sampled_data.reset_index(drop=True),这里多打了一个l,导致重置索引的操作完全没作用到后续使用的sampled_data上,原DataFrame的索引问题没解决,才触发了越界错误。
修正后的函数代码:
def preprocess_data(sampled_data): # Reset index to ensure uniqueness sampled_data = sampled_data.reset_index(drop=True) # Pivot the dataframe to create user-item interaction matrix user_item_matrix = sampled_data.pivot(index='srch_id', columns='prop_id', values='booking_bool').fillna(0) return user_item_matrix user_item_matrix = preprocess_data(sampled_data)
针对大型DataFrame的分块处理,补充两个注意点:
- 分块处理时,每一块都要单独重置索引,不要依赖合并后的整体索引,避免出现索引不连续或超出范围的情况。
- 检查
srch_id是否存在重复值,如果同一个srch_id对应多条记录,pivot可能产生歧义,建议用pivot_table指定聚合方式处理重复的用户-物品组合:# 用pivot_table替代pivot,处理重复组合 user_item_matrix = sampled_data.pivot_table(index='srch_id', columns='prop_id', values='booking_bool', aggfunc='max').fillna(0)
内容的提问来源于stack exchange,提问作者MarsLos10
相关产品推荐
相关产品推荐

