如何创建pandas.dtype常量,用于合成数据类型数组处理空列?
创建pandas.dtype常量的几种方法
1. 直接通过pd.dtype()构造指定类型
传入类型字符串即可生成对应的pandas.dtype常量,适用于常规数据类型:
import pandas as pd # 常规整数类型 int64_dtype = pd.dtype('int64') # 常规浮点类型 float64_dtype = pd.dtype('float64') # pandas专用字符串类型(推荐替代object类型) string_dtype = pd.dtype('string') # 类别类型 category_dtype = pd.dtype('category')
2. 实例化pandas内置的可空类型类
如果需要支持空值(NaN)的数值类型,直接使用pandas提供的大写开头的可空类型类实例化:
# 可空整数类型(允许列中存在NaN) nullable_int_dtype = pd.Int64Dtype() # 可空浮点类型 nullable_float_dtype = pd.Float64Dtype() # 可空布尔类型 nullable_bool_dtype = pd.BooleanDtype()
3. 组合成数据类型数组用于场景化处理
将上述常量整理成列表,就能根据不同场景(数值列、文本列等)批量处理空值/空列:
# 数值类型集合(含常规与可空类型) numeric_dtypes = [ pd.dtype('int64'), pd.dtype('float64'), pd.Int64Dtype(), pd.Float64Dtype() ] # 文本与类别类型集合 text_category_dtypes = [ pd.dtype('string'), pd.dtype('category') ] # 示例:根据类型批量处理空值 def handle_nulls(df): for col in df.columns: if df[col].dtype in numeric_dtypes: df[col] = df[col].fillna(0) # 数值列填充0 elif df[col].dtype in text_category_dtypes: df[col] = df[col].fillna('') # 文本列填充空字符串 return df
内容的提问来源于stack exchange,提问作者WestCoastProjects
相关产品推荐
相关产品推荐

