从含1000列的大型Pandas DataFrame中筛选非指定前缀列
筛选Pandas DataFrame中不以特定前缀开头的列
你需要筛选列名不以指定前缀列表中任意字符串开头的列,生成新的DataFrame。先说明下你给的示例代码逻辑不对——它是从前缀列表里找不在DataFrame列名里的元素,和你要的筛选列的需求不匹配,下面给你两种正确的实现方式:
方法一:列表推导式(直观易懂)
先遍历所有列名,筛选出不以任何指定前缀开头的列,再用这些列索引原始DataFrame:
import pandas as pd # 假设df是你的原始大型DataFrame prefixes = ['fit test and human sorry',"perspectum liver scanaccurate","the following tests were not","apo b to a ratio.this is a useful","psa"] # 统一转小写避免大小写干扰,筛选符合条件的列 selected_columns = [col for col in df.columns if not col.lower().startswith(tuple(p.lower() for p in prefixes))] df_new = df[selected_columns]
方法二:用filter+正则表达式(更简洁)
利用Pandas的filter方法结合正则负向前瞻,直接匹配符合要求的列:
import re import pandas as pd prefixes = ['fit test and human sorry',"perspectum liver scanaccurate","the following tests were not","apo b to a ratio.this is a useful","psa"] # 转义前缀中的特殊字符(比如点号),拼接成正则规则 regex_pattern = r'^(?!' + '|'.join(re.escape(p.lower()) for p in prefixes) + ')' df_new = df.filter(regex=regex_pattern, axis=1)
两种方法都能高效处理1000列的DataFrame,推荐第一种,逻辑更直观,不容易出错。
内容的提问来源于stack exchange,提问作者WhoamI
相关产品推荐
相关产品推荐

