如何一次性检测Pandas DataFrame多列是否存在混合数据类型
Pandas批量检测混合数据类型列实现方案
核心实现代码
针对200万行、15列的DataFrame场景,以下代码可直接输出所有存在混合数据类型的列名:
import pandas as pd # 筛选所有object类型列,非object类型列不存在混合基础数据类型问题 object_cols = df_orders.select_dtypes(include=['object']).columns mixed_type_cols = [] for col in object_cols: # 统计列内所有元素的类型数量,大于1即存在混合类型 if df_orders[col].apply(type).nunique() > 1: mixed_type_cols.append(col) # 直接输出结果 print("存在混合数据类型的列:", mixed_type_cols)
如果偏好更简洁的写法,也可以用列表推导式实现:
mixed_type_cols = [col for col in df_orders.columns if df_orders[col].apply(type).nunique() > 1] print(mixed_type_cols)
常见场景优化
如果你的列中存在空值(NaN),空值本身类型为float,可能会导致非混合类型的列被误判,可使用以下版本忽略空值的影响:
object_cols = df_orders.select_dtypes(include=['object']).columns mixed_type_cols = [] for col in object_cols: # 先剔除空值再统计类型数量 unique_type_cnt = df_orders[col].dropna().apply(type).nunique() if unique_type_cnt > 1: mixed_type_cols.append(col) print("存在混合数据类型的列:", mixed_type_cols)
性能说明
15列的规模下,即使是200万行数据,全量遍历的耗时也在可接受范围内,无需额外采样避免漏判。
内容的提问来源于stack exchange,提问作者Rajiv Jani
相关产品推荐
相关产品推荐

