Pandas中用map执行strip()后NaN未变更却显示不等?
为什么DataFrame处理前后的NaN值看起来一致但比较结果为False?
我有一个包含多种数据类型(文本、整数、浮点数、时间等)的DataFrame,为使后续代码正常运行,尝试用map方法对字符串类型条目执行strip()去除首尾空格。但运行代码后发现,处理前后的DataFrame中NaN条目看起来完全一致,但直接比较时却显示不等。
测试代码
import pandas as pd import numpy as np df1 = pd.DataFrame(np.array(([np.nan, 2, 3], [4, 5, 6])), columns=["one", "two", "three"]) print(df1) print("") df2 = df1.map(lambda x: x.strip() if isinstance(x, str) else x) print(df2) print("") print(df1==df2) print("") cell1 = df1.at[0, "one"] cell2 = df2.at[0, "one"] print(cell1, type(cell1)) print(cell2, type(cell2)) print(cell1==cell2)
运行输出
one two three 0 NaN 2.0 3.0 1 4.0 5.0 6.0 one two three 0 NaN 2.0 3.0 1 4.0 5.0 6.0 one two three 0 False True True 1 True True True nan <class 'numpy.float64'> nan <class 'numpy.float64'> False
原因分析与解决方法
核心问题出在NaN的特性上:按照IEEE 754浮点数标准,NaN(非数值)是唯一一个和自身不相等的值,这是硬性规定,和你对DataFrame的处理逻辑无关。
你的map逻辑里,因为NaN不属于字符串类型,所以直接返回了原值,df1和df2里的NaN确实是同一个值,但用==直接比较时,依然会返回False。
如果要正确判断包含NaN的DataFrame是否相等,推荐用这两种方式:
- 使用pandas内置的
equals()方法,它会自动处理NaN的相等性判断:print(df1.equals(df2)) # 输出True - 单独判断某个值是否为NaN时,用
np.isnan():print(np.isnan(cell1) and np.isnan(cell2)) # 输出True
内容的提问来源于stack exchange,提问作者python3 programmer
相关产品推荐
相关产品推荐

