You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除CSV中无法解析为数值的非数字值(通用方法)

问题场景

我有一个大型数据表,示例如下:
CSV文件示例

我想要删除CSV中的所有字符串值,曾尝试手动枚举列名调用drop方法实现:

df.drop(['document.children.children.id', 'document.id', 'document.name', 'document.type', 'document.children.name', 'document.children.type', 'document.children.children.name', 'document.children.children.type', 'document.children.children.blendMode', 'document.children.children.children.blendMode', 'document.children.children.children.fills.blendMode', 'document.children.children.children.fills.type'], axis=1, inplace=True )

但处理其他结构的设计数据时,手动指定列名的方法完全无法复用,需要通用化的实现方案。

解决方案

不需要手动枚举列名,基于pandas的类型判断接口即可实现适配任意表结构的处理逻辑。

  • 基础场景:整列均为字符串值的批量删除
    直接用select_dtypes方法按列类型筛选,排除所有字符串类型的列即可,代码可以适配任意列名、任意结构的DataFrame:

    # 自动丢弃所有存储类型为object、string的字符串列,保留数值、布尔等非字符串列
    df = df.select_dtypes(exclude=['object', 'string'])
    

    说明:pandas默认将字符串列存储为object类型,若手动开启了StringDtype扩展类型则会标记为string,两种类型同时排除即可覆盖所有常规纯字符串列。

  • 混合类型场景:列内同时存在数值和字符串的处理
    如果部分列混存了数值和字符串,pandas会将整列识别为object类型,直接筛选会误保留列内的字符串内容,可以先做强制类型转换再清理:

    import numpy as np
    
    # 遍历所有列,尝试将内容转为数值,无法转换的字符串统一转为空值NaN
    for col in df.columns:
        df[col] = pd.to_numeric(df[col], errors='coerce')
    
    # 转换后纯字符串列会变为全空列,直接删除全空列即可
    df = df.dropna(axis=1, how='all')
    
  • 单元格级清理:不删整列仅清除单元格内的字符串
    如果需要保留列结构,仅把单元格里的字符串值替换为空、保留同列的有效数值,可以用逐元素判断的方式处理:

    import numpy as np
    
    df = df.applymap(lambda x: np.nan if isinstance(x, str) else x)
    

内容的提问来源于stack exchange,提问作者Sayuru De Alwis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 04:48:34