You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中精准排除指定列问题:误删关联列如何修复?

问题描述

我的数据包含格式为ABC_number_AX和ABC_number_AX_MED的列,希望仅排除ABC_number_AX列。编写的正则过滤代码如下:

# Define patterns to filter out patterns 
patterns = ['_AX','_B2']

# Create a regular expression pattern to match any of the defined patterns
pattern = '|'.join(map(re.escape, patterns))

# List comprehension to filter out columns based on the pattern
filtered_columns = [col for col in df.columns if not re.search(pattern, col)]

# Create a new DataFrame with filtered columns
df= df[filtered_columns]

但该代码同时误删了ABC_number_AX_MED列,尝试将模式改为_AX$后仍未得到预期结果,请问该如何修复?


修复方案

核心问题拆解

  1. 最初用_AX匹配时,ABC_number_AX_MED包含_AX子串,所以被误过滤;
  2. _AX$没生效,大概率是列名末尾存在隐形空格,或者正则匹配逻辑未考虑边界情况。

可行解决方法

方法一:精准匹配结尾(处理空格)

针对列名可能带末尾空格的情况,用正则匹配_AX结尾并忽略后续空格,同时排除_B2结尾的列:

import re

# 正则匹配:以_AX结尾(允许末尾空格),或者以_B2结尾
pattern = r'(_AX\s*$)|(_B2\s*$)'

# 过滤逻辑:保留不匹配上述模式的列
filtered_columns = [col for col in df.columns if not re.search(pattern, col)]
df = df[filtered_columns]

方法二:直接列规则筛选(更直观)

跳过正则,用字符串方法直接筛选要排除的列:

# 筛选所有以_AX结尾,但不是_AX_MED结尾的列,再加上_B2结尾的列
exclude_cols = [
    col for col in df.columns 
    if (col.rstrip().endswith('_AX') and not col.rstrip().endswith('_AX_MED')) 
    or col.rstrip().endswith('_B2')
]

# 删除目标列
df = df.drop(columns=exclude_cols)

方法三:先清理列名再过滤

如果确认是列名的空格导致匹配失效,先统一清理列名:

# 去除所有列名的首尾空格
df.columns = [col.strip() for col in df.columns]

# 再用简单的结尾匹配过滤
filtered_columns = [
    col for col in df.columns 
    if not col.endswith('_AX') and not col.endswith('_B2')
]
df = df[filtered_columns]

快速排查技巧

先打印所有列名,检查是否有空格或特殊字符干扰:

print(df.columns.tolist())

内容的提问来源于stack exchange,提问作者P.Chakytei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 06:03:21