You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas:如何删除不在指定索引集合中的数据行

解决CSV文件中删除指定范围外数据行的问题

需求:删除CSV文件中name列值不在dirs列表中的数据行,dirs列表由目录扫描生成。


现有代码与问题重现

生成目标列表

import os
dirs = os.listdir('foo/')
dirs
# 输出: ['boo', 'aoo', 'coo', 'doo']

读取CSV数据

import pandas as pd
file = pd.read_csv('tada.csv')
file.head()
# 输出:
#     name  height  weight gender
# 0   aoo     212     253      M
# 1   boo     175     243      M
# 2   coo     190     244      M
# 3   doo     162     288      F
# 4   too     222     240      M

错误尝试与结果

你先筛选出了需要保留的行索引:

index = []
idx = []
for dname in dirs:
    a = file.index[file['name'] == dname].tolist()
    index.append(a)

for i in index:
    for j in i:
        idx.append(j)
print(idx)
# 输出: [1, 0, 2, 3]

但错误地删除了这些需要保留的索引:

for i in idx:
    file.drop(i,axis=0,inplace=True)
print(file)
# 输出:
#   name  height  weight gender
# 4  too     222      240      M
# 5  yoo     272     230      F
# 6  poo     200      23      F

错误原因

你搞反了操作逻辑:idx是需要保留的行索引,但你却循环删除这些索引,最终留下的自然是不需要的数据。此外,循环调用drop效率极低,还可能因为DataFrame实时变化导致索引对应错误。


正确实现方法

方法一:直接筛选保留符合条件的行(推荐)

通过布尔索引直接筛选出name在dirs中的行,这是Pandas最简洁高效的方式:

# 筛选并保留目标行,重新赋值给原变量
file = file[file['name'].isin(dirs)]
print(file)
# 输出:
#     name  height  weight gender
# 0   aoo     212     253      M
# 1   boo     175     243      M
# 2   coo     190     244      M
# 3   doo     162     288      F

方法二:批量删除目标索引

如果一定要使用drop,先找出所有需要删除的索引,一次性批量操作:

# 找出name不在dirs中的行的索引
to_drop = file.index[~file['name'].isin(dirs)]
# 批量删除这些索引
file.drop(to_drop, axis=0, inplace=True)
print(file)
# 输出与方法一一致

内容的提问来源于stack exchange,提问作者jabbar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 13:35:34