You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

构建推荐系统时持续触发IndexError索引越界错误,求排查解决

解决IndexError问题的步骤

首先看你代码里的明显错误:

  • 重置索引的那行变量名拼写错误:sampld_data = sampled_data.reset_index(drop=True),这里多打了一个l,导致重置索引的操作完全没作用到后续使用的sampled_data上,原DataFrame的索引问题没解决,才触发了越界错误。

修正后的函数代码:

def preprocess_data(sampled_data):
    # Reset index to ensure uniqueness
    sampled_data = sampled_data.reset_index(drop=True)
    
    # Pivot the dataframe to create user-item interaction matrix
    user_item_matrix = sampled_data.pivot(index='srch_id', columns='prop_id', values='booking_bool').fillna(0)
    
    return user_item_matrix

user_item_matrix = preprocess_data(sampled_data)

针对大型DataFrame的分块处理,补充两个注意点:

  • 分块处理时,每一块都要单独重置索引,不要依赖合并后的整体索引,避免出现索引不连续或超出范围的情况。
  • 检查srch_id是否存在重复值,如果同一个srch_id对应多条记录,pivot可能产生歧义,建议用pivot_table指定聚合方式处理重复的用户-物品组合:
    # 用pivot_table替代pivot,处理重复组合
    user_item_matrix = sampled_data.pivot_table(index='srch_id', columns='prop_id', values='booking_bool', aggfunc='max').fillna(0)
    

内容的提问来源于stack exchange,提问作者MarsLos10

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 18:02:06