如何获取Pandas数据框各列的真实数据类型?
获取Pandas数据列的真实数据类型
针对你遇到的Pandas将字符串、含NaN的布尔列标记为object类型的问题,以下是几种实用的解决方法:
1. 使用pd.api.types.infer_dtype()精准推断类型
这个函数能绕过object的外壳,返回列的实际逻辑类型,比如字符串列会返回'string',含NaN的布尔列会返回'boolean'。
示例代码:
import pandas as pd df = pd.DataFrame({ 'str_col': ['a', 'b', 'c'], 'bool_with_nan': [True, False, None], 'mixed_obj': [1, 'a', True] }) for col in df.columns: real_dtype = pd.api.types.infer_dtype(df[col], skipna=True) print(f"列 {col} 的真实类型: {real_dtype}")
输出:
列 str_col 的真实类型: string 列 bool_with_nan 的真实类型: boolean 列 mixed_obj 的真实类型: mixed-integer
2. 转换为Pandas专属的可空类型(推荐)
Pandas 1.0+版本支持专属的可空数据类型,转换后能直接在df.info()或df.dtypes中显示真实类型:
- 字符串列:转换为
StringDtype - 含NaN的布尔列:转换为
BooleanDtype
示例代码:
# 转换字符串列 df['str_col'] = df['str_col'].astype('string') # 转换可空布尔列 df['bool_with_nan'] = df['bool_with_nan'].astype('boolean') print(df.dtypes)
输出:
str_col string[python] bool_with_nan boolean mixed_obj object dtype: object
3. 检查列内元素的类型分布
如果列是混合类型,可以通过统计元素类型来判断真实情况:
def get_most_common_dtype(series): type_counts = series.apply(type).value_counts() return type_counts.index[0].__name__ if not type_counts.empty else None for col in df.columns: print(f"列 {col} 的最常见类型: {get_most_common_dtype(df[col])}")
注意:NaN的类型是float,处理含NaN的布尔列时,建议结合skipna=True过滤NaN后再判断。
4. 针对布尔列的特殊判断
如果怀疑某列是含NaN的布尔列,可以用以下方式验证:
def is_boolean_with_nan(series): # 过滤NaN后检查剩余元素是否都是布尔值 non_na = series.dropna() return non_na.apply(lambda x: isinstance(x, bool)).all() for col in df.columns: if is_boolean_with_nan(df[col]): print(f"列 {col} 是含NaN的布尔列")
内容的提问来源于stack exchange,提问作者Marcus Anthony
相关产品推荐
相关产品推荐

