运行旅行推荐系统代码遇IndexError:索引越界问题求助
问题排查:智能旅行推荐系统IndexError报错解决
问题背景
运行智能旅行推荐系统代码时,触发IndexError: indices are out-of-bounds和IndexError: positional indexers are out-of-bounds错误,无法确定是包更新还是循环逻辑问题,相关代码及报错栈如下:
核心代码片段
%%capture final = dict() final['timeofday'] = [] final['image'] = [] final['name'] = [] final['location'] = [] final['price'] = [] final['rating'] = [] final['category'] = [] for i in range(1,(end_date - begin_date).days+2): for j in range(2): final['timeofday'].append('Morning') for j in range(2): final['timeofday'].append('Evening') for i in range(len(final['timeofday'])): if i%4 == 0: final = top_recc(with_url, final) else: final = find_closest(with_url, final['location'][-1],final['timeofday'][i], final)
完整报错栈
IndexError Traceback (most recent call last) File ~\anaconda3\lib\site-packages\pandas\core\indexing.py:1482, in _iLocIndexer._get_list_axis(self, key, axis) 1481 try: -> 1482 return self.obj._take_with_is_copy(key, axis=axis) 1483 except IndexError as err: 1484 # re-raise with different error message File ~\anaconda3\lib\site-packages\pandas\core\generic.py:3716, in NDFrame._take_with_is_copy(self, indices, axis) 3709 """ 3710 Internal version of the `take` method that sets the `_is_copy` 3711 attribute to keep track of the parent dataframe (using in indexing (...) 3714 See the docstring of `take` for full explanation of the parameters. 3715 """ -> 3716 result = self.take(indices=indices, axis=axis) 3717 # Maybe set copy if we didn't actually change the index. File ~\anaconda3\lib\site-packages\pandas\core\generic.py:3703, in NDFrame.take(self, indices, axis, is_copy, **kwargs) 3701 self._consolidate_inplace() -> 3703 new_data = self._mgr.take( 3704 indices, axis=self._get_block_manager_axis(axis), verify=True 3705 ) 3706 return self._constructor(new_data).__finalize__(self, method="take") File ~\anaconda3\lib\site-packages\pandas\core\internals\managers.py:897, in BaseBlockManager.take(self, indexer, axis, verify) 896 n = self.shape[axis] -> 897 indexer = maybe_convert_indices(indexer, n, verify=verify) 899 new_labels = self.axes[axis].take(indexer) File ~\anaconda3\lib\site-packages\pandas\core\indexers\utils.py:292, in maybe_convert_indices(indices, n, verify) 291 if mask.any(): -> 292 raise IndexError("indices are out-of-bounds") 293 return indices IndexError: indices are out-of-bounds The above exception was the direct cause of the following exception: IndexError Traceback (most recent call last) Input In [8], in <cell line: 16>() 16 for i in range(len(final['timeofday'])): 17 if i%4 == 0: ---> 18 final = top_recc(with_url, final) 19 else: 20 final = find_closest(with_url, final['location'][-1],final['timeofday'][i], final) File ~\Desktop\Intelligent-Travel-Recommendation-System-master\attractions_recc.py:114, in top_recc(with_url, final) 112 i=0 113 while(1): -> 114 first_recc = with_url.iloc[[i]] 115 if(first_recc['name'].values.T[0] not in final['name']): 116 final['name'].append(first_recc['name'].values.T[0]) File ~\anaconda3\lib\site-packages\pandas\core\indexing.py:967, in _LocationIndexer.__getitem__(self, key) 964 axis = self.axis or 0 966 maybe_callable = com.apply_if_callable(key, self.obj) -> 967 return self._getitem_axis(maybe_callable, axis=axis) File ~\anaconda3\lib\site-packages\pandas\core\indexing.py:1511, in _iLocIndexer._getitem_axis(self, key, axis) 1509 # a list of integers 1510 elif is_list_like_indexer(key): -> 1511 return self._get_list_axis(key, axis=axis) 1513 # a single integer 1514 else: 1515 key = item_from_zerodim(key) File ~\anaconda3\lib\site-packages\pandas\core\indexing.py:1485, in _iLocIndexer._get_list_axis(self, key, axis) 1482 return self.obj._take_with_is_copy(key, axis=axis) 1483 except IndexError as err: 1484 # re-raise with different error message -> 1485 raise IndexError("positional indexers are out-of-bounds") from err IndexError: positional indexers are out-of-bounds
报错原因分析
- 核心触发点:错误来自
top_recc函数的第114行with_url.iloc[[i]]——当变量i的取值超过with_url这个DataFrame的最大行索引时,就会触发索引越界。 - 逻辑缺陷:
top_recc里的while(1)是无限循环,没有设置终止条件,每次循环i递增后最终会超出with_url的行数上限。- 外层循环的次数由
len(final['timeofday'])决定,这个值根据旅行天数生成(每天4个时段),但如果旅行天数对应的推荐需求超过了with_url里的景点总数,就会导致反复调用top_recc时耗尽所有可推荐景点,最终触发索引越界。
- 排除包更新因素:pandas中
iloc的索引逻辑是基础行为,版本更新不会改变“索引超出范围报错”的规则,因此不是包更新导致的问题。
解决方案
1. 修复top_recc函数的循环逻辑
给无限循环加上终止条件,避免索引越界:
def top_recc(with_url, final): i = 0 max_valid_index = len(with_url) - 1 # 获取DataFrame的最大合法索引 while i <= max_valid_index: first_recc = with_url.iloc[[i]] target_name = first_recc['name'].values.T[0] if target_name not in final['name']: # 补充其他字段的赋值逻辑 final['name'].append(target_name) final['image'].append(first_recc['image'].values.T[0]) final['location'].append(first_recc['location'].values.T[0]) final['price'].append(first_recc['price'].values.T[0]) final['rating'].append(first_recc['rating'].values.T[0]) final['category'].append(first_recc['category'].values.T[0]) return final i += 1 # 所有景点已推荐完毕,直接返回原final return final
2. 限制外层循环的执行次数
在外层循环中加入检查,当没有更多可推荐的景点时,提前终止循环:
max_available_recc = len(with_url) # 可推荐的最大景点数 for i in range(len(final['timeofday'])): # 若已推荐完所有景点,直接终止循环 if len(final['name']) >= max_available_recc: break if i%4 == 0: final = top_recc(with_url, final) else: final = find_closest(with_url, final['location'][-1],final['timeofday'][i], final)
3. 可选优化:提升重复检查效率
如果final['name']的长度很大,target_name not in final['name']的效率很低,可以改用集合来记录已推荐的景点:
# 初始化时添加集合存储已推荐名称 final['recommended_names'] = set() # 在top_recc里修改检查逻辑 if target_name not in final['recommended_names']: final['name'].append(target_name) final['recommended_names'].add(target_name) # 其他字段赋值...
内容的提问来源于stack exchange,提问作者Mahdi Karblaee
相关产品推荐
相关产品推荐

