Pandas:如何删除不在指定索引集合中的数据行
解决CSV文件中删除指定范围外数据行的问题
需求:删除CSV文件中name列值不在dirs列表中的数据行,dirs列表由目录扫描生成。
现有代码与问题重现
生成目标列表
import os dirs = os.listdir('foo/') dirs # 输出: ['boo', 'aoo', 'coo', 'doo']
读取CSV数据
import pandas as pd file = pd.read_csv('tada.csv') file.head() # 输出: # name height weight gender # 0 aoo 212 253 M # 1 boo 175 243 M # 2 coo 190 244 M # 3 doo 162 288 F # 4 too 222 240 M
错误尝试与结果
你先筛选出了需要保留的行索引:
index = [] idx = [] for dname in dirs: a = file.index[file['name'] == dname].tolist() index.append(a) for i in index: for j in i: idx.append(j) print(idx) # 输出: [1, 0, 2, 3]
但错误地删除了这些需要保留的索引:
for i in idx: file.drop(i,axis=0,inplace=True) print(file) # 输出: # name height weight gender # 4 too 222 240 M # 5 yoo 272 230 F # 6 poo 200 23 F
错误原因
你搞反了操作逻辑:idx是需要保留的行索引,但你却循环删除这些索引,最终留下的自然是不需要的数据。此外,循环调用drop效率极低,还可能因为DataFrame实时变化导致索引对应错误。
正确实现方法
方法一:直接筛选保留符合条件的行(推荐)
通过布尔索引直接筛选出name在dirs中的行,这是Pandas最简洁高效的方式:
# 筛选并保留目标行,重新赋值给原变量 file = file[file['name'].isin(dirs)] print(file) # 输出: # name height weight gender # 0 aoo 212 253 M # 1 boo 175 243 M # 2 coo 190 244 M # 3 doo 162 288 F
方法二:批量删除目标索引
如果一定要使用drop,先找出所有需要删除的索引,一次性批量操作:
# 找出name不在dirs中的行的索引 to_drop = file.index[~file['name'].isin(dirs)] # 批量删除这些索引 file.drop(to_drop, axis=0, inplace=True) print(file) # 输出与方法一一致
内容的提问来源于stack exchange,提问作者jabbar
相关产品推荐
相关产品推荐

