You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中先移除列内数值再统计非数值型值的频次?

快速统计DataFrame字符串列中的非数值型值频次

针对100万行的大数据量,推荐使用矢量化操作实现高效筛选,避免逐行处理的性能损耗。具体方案如下:

核心思路

利用pd.to_numeric的errors='coerce'参数,将可转换为数值的字符串转为对应数值,无法转换的非数值型字符串转为NaN;随后通过判断NaN筛选出非数值型数据,最后统计频次。

代码实现

import pandas as pd

df = pd.DataFrame({'col1':['1','2','C','T','A','00400'],
                   'col2':['3241','H2','C8','T4','123','0000']})

# 生成非数值型数据的布尔掩码
non_numeric_mask = pd.to_numeric(df['col1'], errors='coerce').isna()

# 筛选非数值型数据并统计频次
non_numeric_counts = df['col1'][non_numeric_mask].value_counts()

print(non_numeric_counts)

输出结果

A    1
C    1
T    1
Name: col1, dtype: int64

性能说明

这种方式依赖Pandas的矢量化运算,处理100万行数据的速度远快于apply逐行判断的方法,能有效避免性能瓶颈。

多列批量处理(可选)

如果需要对所有字符串列批量统计,可循环处理:

for col in df.columns:
    non_numeric_mask = pd.to_numeric(df[col], errors='coerce').isna()
    counts = df[col][non_numeric_mask].value_counts()
    print(f"列 {col} 的非数值型频次统计:")
    print(counts)
    print("---")

内容的提问来源于stack exchange,提问作者frank

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 02:02:40