pandas iloc报single positional indexer is out-of-bounds错误
报错原因
报错核心是两个逻辑问题,外加一个pandas索引使用错误:
- 你初始化的
dfTest是只有列名、0行数据的空DataFrame,iloc是按行的位置序号做索引,要求传入的序号必须小于当前DataFrame的总行数。空表总行数是0,你第一次循环取iloc[0]就会触发越界,直接抛出single positional indexer is out-of-bounds错误。 - 就算你提前给DataFrame分配了足够行数,现有逻辑也存不对数据:内层循环的
enumerate是按每个子文件夹单独从0开始计数的,每遍历一个新的子文件夹,k都会重置为0,会直接覆盖之前写入的同序号行内容,最终只会保留最后一个子文件夹的文件列表。 - 你写的
dfTest.iloc[k]['filename']= j属于pandas里的链式索引,这种写法本质是先取切片副本再给副本赋值,就算索引不越界,值也大概率不会写入原DataFrame,还会触发SettingWithCopyWarning警告。
修复方法
最稳妥高效的写法是先遍历收集所有文件名到列表,再一次性构造DataFrame,不要在循环里逐行做iloc/loc赋值,既慢又容易出索引问题,另外路径拼接用os.path.join代替手动拼/,避免跨平台或者多斜杠的路径错误。
修正后代码:
import os import pandas as pd # 定义根目录路径 test_root = '/home/ubuntu/imageTrain_dobby/SKJEWELLERY/BC4U/google_version/v1.1/lingyau_lee/output/test' file_list = [] # 遍历根目录下的子文件夹 for category in os.listdir(test_root): category_path = os.path.join(test_root, category) # 跳过非目录的文件,避免遍历报错 if not os.path.isdir(category_path): continue # 遍历子文件夹下的所有文件 for filename in os.listdir(category_path): # 如果需要存文件全路径就传入os.path.join(category_path, filename),只存文件名就传filename file_list.append({ 'filename': filename, 'class': category # 子文件夹名可直接作为类别值,匹配初始定义的class列 }) # 一次性构造DataFrame dfTest = pd.DataFrame(file_list)
如果你确实需要在循环中逐行追加数据(不推荐,数据量大时性能很差),可以用loc按长度索引追加新行,写法如下:
# 初始化空表 dfTest = pd.DataFrame(columns=['filename', 'class']) test_root = '/home/ubuntu/imageTrain_dobby/SKJEWELLERY/BC4U/google_version/v1.1/lingyau_lee/output/test' for category in os.listdir(test_root): category_path = os.path.join(test_root, category) if not os.path.isdir(category_path): continue for filename in os.listdir(category_path): # 逐行追加,新行的索引为当前表的总行数 dfTest.loc[len(dfTest)] = [filename, category]
内容的提问来源于stack exchange,提问作者lingyau lee
相关产品推荐
相关产品推荐

