Python中如何使用索引列表批量删除列表或DataFrame中的对应元素
问题解决方法
方法1:使用drop方法批量删除指定索引
pandas的DataFrame.drop()原生支持传入索引标签列表批量删除对应行,完全不会出现索引偏移问题,该方法是按索引标签匹配删除,不是按行位置删除,刚好匹配你的需求:
# 假设你已获取的测试集对应原DF索引列表为 test_index_list test_set = df.loc[test_index_list].copy() # 提取测试集 train_set = df.drop(test_index_list).copy() # 批量删除测试集行得到训练集
方法2:通过索引差集运算拆分
你提到的差集运算完全可以实现,pandas的索引对象内置difference()差集计算方法:
# 计算排除测试集索引后的训练集索引 train_index_list = df.index.difference(test_index_list) train_set = df.loc[train_index_list].copy() test_set = df.loc[test_index_list].copy()
该方法更适配存在重复索引、需要提前校验索引正确性的场景。
补充说明
- 不要使用for循环配合
pop()逐行删除,该操作不仅执行效率极低,还会因为逐行删除后行位置动态变化,导致后续的位置索引匹配错误,就是你遇到的偏移问题。 - 如果你还未生成测试集的随机索引,可直接使用
sklearn.model_selection.train_test_split()完成自动拆分,无需手动生成随机索引:
from sklearn.model_selection import train_test_split # test_size=0.2 表示拆分20%数据为测试集,random_state固定随机种子保证结果可复现 train_set, test_set = train_test_split(df, test_size=0.2, random_state=42)
内容的提问来源于stack exchange,提问作者Regina Briseño
相关产品推荐
相关产品推荐

