如何在Python中用np.full创建NaN或Null值数组?
在Python中创建含缺失值的数组(对应R的NA向量)
核心问题:字符串类型的缺失值陷阱
在R里字符型向量能直接赋值NA,但Python中str dtype的numpy数组无法存储真正的缺失值——numpy的字符串数组本质是固定长度的字节序列,会把np.nan转成字符串'nan',把None转成'None',这两类字符串都不会被pd.isna识别为缺失值,这就是你测试失败的原因。
正确实现方式
要让pd.isna能识别缺失值,得用支持缺失值的类型,下面提供两种可行方案:
方法1:用numpy的object dtype存储None
object dtype可以存储Python原生对象,None会被pd.isna正确识别为缺失值:
import numpy as np import pandas as pd import random # 创建5个None的数组,指定dtype为object unknowns = np.full(shape=5, fill_value=None, dtype='object') # 测试代码 random.seed(1177) categories = np.random.choice(['web', 'software', 'hardware', 'biotech'], size=15, replace=True) categories = np.concatenate([categories, unknowns]) example = pd.DataFrame(data={'categories': categories}) example['transformed'] = [x if not pd.isna(x) else 'unknown' for x in example['categories']] print(example['transformed'].value_counts())
运行后会看到unknown的计数为5,符合预期。
方法2:用pandas专用字符串类型更规范
如果直接基于pandas处理,推荐使用StringDtype,它原生支持pd.NA缺失值:
import pandas as pd import random # 创建含5个pd.NA的Series,指定dtype为string unknowns = pd.Series([pd.NA]*5, dtype='string') # 测试代码 random.seed(1177) categories = pd.Series(np.random.choice(['web', 'software', 'hardware', 'biotech'], size=15, replace=True), dtype='string') categories = pd.concat([categories, unknowns], ignore_index=True) example = pd.DataFrame(data={'categories': categories}) example['transformed'] = example['categories'].fillna('unknown') print(example['transformed'].value_counts())
这种方式更贴合pandas生态,避免numpy字符串类型的兼容性问题。
不同变量类型的实现差异
- 数值型数组:可直接用
np.full(shape=5, fill_value=np.nan, dtype='float'),np.nan作为数值型缺失值,能被pd.isna识别。 - 布尔型数组:numpy不支持布尔型缺失值,需转成
objectdtype存储None,或用pandas的BooleanDtype存储pd.NA。 - 字符串/对象型数组:只有存储
None(object dtype)或pd.NA(pandas StringDtype)才会被pd.isna识别为缺失值;numpy的str/Udtype会把缺失值转成字符串字面量,无法被识别。
内容的提问来源于stack exchange,提问作者Pearl
相关产品推荐
相关产品推荐

