pandas如何基于多个条件批量筛选数据集指定列
问题原因
你的代码中存在两处可优化/错误点:
str.contains('gibberish','tech')写法不符合pandas语法:Series.str.contains()的第二个参数不是第二个匹配字符串,而是大小写控制、正则开关等配置参数,因此该语句仅过滤了包含gibberish的列,未过滤tech列endswith方法原生支持传入多后缀元组,无需拆分写多段过滤逻辑,可简化代码
正确代码
针对上百个变量的维护需求,推荐将过滤规则拆分归类,后续新增过滤规则直接修改对应列表/元组即可:
import pandas as pd # 定义排除规则,后续新增规则直接修改对应位置即可 exclude_suffix = ('_pct', '_ln') exclude_cols = ['gibberish', 'tech'] df_selected = my_df.loc[:, ~my_df.columns.str.endswith(exclude_suffix) & ~my_df.columns.isin(exclude_cols) ]
输出结果
运行上述代码后得到预期结果:
| id | type | sales | sales_roi | flag |
|---|---|---|---|---|
| 1 | corp | 34567 | 0.10 | 0 |
| 2 | smb | 2190 | 0.21 | 1 |
| 3 | smb | 1870 | 0.22 | 0 |
| 4 | corp | 22000 | 0.15 | 1 |
| 5 | mid | 10000 | 0.16 | 1 |
内容的提问来源于stack exchange,提问作者Alexis
相关产品推荐
相关产品推荐

