You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中筛选str、float、int标量列并排除对象列

筛选DataFrame中仅包含int、float、字符串的列(排除字典/对象列)

你遇到的问题是:df.select_dtypes(include=[int, float, str])会触发TypeError,因为pandas不接受str作为dtype参数;而include=[int, float, object]会把存储字典、类实例等的object列也纳入,无法区分纯字符串列。

解决方案

我们可以分两步筛选目标列:

  1. 先提取int和float类型的列
  2. 对object类型列,检查每一列的元素是否为字符串,筛选出纯字符串列
  3. 合并两类列得到最终结果

代码示例

测试数据构建

import pandas as pd

# 创建混合类型测试DataFrame
df = pd.util.testing.makeMixedDataFrame()
# 添加一列存储字典的对象列
df["E"] = df.A.map(lambda x: {"E": x**2})

筛选逻辑

# 提取数值类型(int、float)列
numeric_cols = df.select_dtypes(include=[int, float]).columns

# 筛选纯字符串类型的object列
string_cols = [
    col for col in df.select_dtypes(include=[object]).columns
    if df[col].apply(lambda elem: isinstance(elem, str)).all()
]

# 合并得到目标列集合,并生成筛选后的DataFrame
target_columns = numeric_cols.union(string_cols)
filtered_df = df[target_columns]

验证结果

print(filtered_df)
# 输出:
#      A     B      C
# 0  0.0  0.00   foo1
# 1  1.0  1.00   foo2
# 2  2.0  2.00   foo3
# 3  3.0  3.00   foo4
# 4  4.0  16.00  foo5

补充说明

  • 如果你的数据中存在少量非字符串元素(如NaN)但主体是字符串,可以将.all()改为.mean() > 0.9这类逻辑来放宽筛选条件
  • 该方法能精准排除存储字典、类实例等非字符串的object列

内容的提问来源于stack exchange,提问作者Janosh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 17:15:43