You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas列中将非浮点数格式的值替换为np.nan

问题:将DataFrame中非浮点数格式的值替换为np.nan

原始数据

import numpy as np
import pandas as pd

d_test = {
    'c1' : ['31', '421', 'sgdsgd', '523.3'],
    'c2' : ['41', np.nan, '412', '412'],
    'test': [1,2,3,4],
}
df_test = pd.DataFrame(d_test)

需求

将c1、c2列内所有非浮点数格式的值替换为np.nan,预期结果:

c1    c2  test
0    31    41     1
1   421   NaN     2
2   NaN   412     3
3  523.3  412     4

尝试的错误代码

df_test[['c1', 'c2']] = df_test[['c1', 'c2']].replace(to_replace=r'^[+-]?([0-9]+([.][0-9]*)?|[.][0-9]+)$', value=np.nan, regex=True)

错误结果

c1   c2  test
0    NaN  NaN     1
1    NaN  NaN     2
2  sgdsgd NaN     3
3    NaN  NaN     4

问题分析与解决方法

你的正则逻辑完全搞反了:当前正则匹配的是符合浮点数格式的字符串,然后把这些合法值替换成了NaN,和需求背道而驰。

方法1:用pd.to_numeric(推荐)

这是最简洁高效的方案,pd.to_numeric的errors='coerce'参数会自动将非数值的字符串转为np.nan,合法数值原样保留:

df_test[['c1', 'c2']] = df_test[['c1', 'c2']].apply(pd.to_numeric, errors='coerce')

方法2:修正正则替换

使用否定前瞻正则,精准匹配不符合浮点数格式的字符串,再替换为np.nan:

df_test[['c1', 'c2']] = df_test[['c1', 'c2']].replace(to_replace=r'^(?![+-]?(\d+(\.\d*)?|\.\d+)$).*$', value=np.nan, regex=True)

解释:(?![+-]?(\d+(\.\d*)?|\.\d+)$)是否定前瞻规则,意思是只要字符串不匹配浮点数格式,就会被选中替换为NaN。


内容的提问来源于stack exchange,提问作者illuminato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 10:40:37