You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas中删除含‘…’的列:已尝试方法无效,求最优方案

问题:移除DataFrame中全为“…”的列

场景说明

有如下表格,其中Samtskhe-Javakheti列的所有值均为“…”:

Kakheti      Tbilisi  Shida Kartli  Kvemo Kartli Samtskhe-Javakheti  
1    447.773080   695.168755    575.344860    492.989720                  …   
2    368.175479   680.922659    449.687764    428.683988                  …   
3     94.253356   381.434387    147.149448    219.399642                  …   
4     38.283124    77.261457     39.516685     29.104063                  …   
5     71.281052     0.720206     74.027796     49.294079                  …   
6      1.463107    15.452695      2.457914      0.000000                  …   

尝试的代码

已尝试以下代码,但dropna和正则匹配列名的方式都无法移除目标列:

import pandas as pd
data =pd.read_excel("https://geostat.ge/media/45425/106_Distribution-of-average-monthly-incomes-per-household-by-regions.xls",
                    skiprows=[0])
data.drop(["Unnamed: 0","Other regions**","Georgia"],axis=1,inplace=True)
data.dropna(axis=0,how='all',inplace=True)
data.dropna(axis=1,how='all',inplace=True)
for column in data.columns:
    if data[column].dtype=="object":
        data[column] =data[column].str.strip()
data = data[data.columns.drop(list(data.filter(regex='…')))]
pd.set_option('display.max_columns', None)
pd.set_option('display.max_rows', 165)
print(data.head(100))

其中重点尝试的移除列代码:

data = data[data.columns.drop(list(data.filter(regex='…')))]

补充说明

发现使用data = data._get_numeric_data()可以得到预期结果(保留所有数值列),但希望找到更合适的方案,预期结果如下:

Kakheti      Tbilisi  Shida Kartli  Kvemo Kartli  Adjara A.R.  
1    447.773080   695.168755    575.344860    492.989720   656.011810   
2    368.175479   680.922659    449.687764    428.683988   576.822184   
3     94.253356   381.434387    147.149448    219.399642   289.083094   
4     38.283124    77.261457     39.516685     29.104063   112.680109   
5     71.281052     0.720206     74.027796     49.294079    17.850210   
6      1.463107    15.452695      2.457914      0.000000     4.489037   

解决方案

方法1:精准过滤全为“…”的列

先确保字符串列的空格已去除(你已完成这一步),再检查每一列是否所有值都是“…”,保留不符合条件的列:

# 确保字符串列去除首尾空格
for column in data.columns:
    if data[column].dtype == "object":
        data[column] = data[column].str.strip()

# 过滤列:保留不全是“…”的列
data = data.loc[:, ~data.apply(lambda col: col.eq("…").all())]

方法2:转数值类型+移除全NaN列

尝试将所有列转为数值类型,无法转换的“…”会被标记为NaN,之后用dropna移除全NaN列,同时还能完成数值列的类型转换:

# 尝试转数值,无效值设为NaN
data = data.apply(pd.to_numeric, errors='coerce')
# 移除全为NaN的列
data = data.dropna(axis=1, how='all')

方法3:用公开API选择数值列(替代内部方法)

_get_numeric_data是pandas内部方法,推荐使用公开稳定的select_dtypes来选择数值类型列:

data = data.select_dtypes(include=['number'])

该方法和_get_numeric_data效果一致,但符合代码规范,更易维护。


内容的提问来源于stack exchange,提问作者data science

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 07:17:03