Python含缺失值的年龄分箱求助:类型转换与分箱报错问题
解决年龄分箱及缺失值处理问题
错误原因分析
- 提前填充字符串导致类型转换失败:先把NaN填充为"unknown"字符串,后续转float时,字符串无法被转换,触发
could not convert string to float: 'N/A'错误。 - 分类操作对象错误:
cat.add_categories仅适用于分类类型(Categorical)数据,你之前的AGE列是数值/字符串混合类型,该操作无效。 - 分箱时类型不兼容:pd.cut无法处理字符串与数值混合的数据,因此出现
not supported between instances of 'int' and 'str'错误。
修正后的代码
import pandas as pd # 1. 清理年龄列:去除%符号,转换为数值,无法转换的标记为NaN Students['AGE'] = pd.to_numeric(Students['AGE'].str.replace('%', '', regex=True), errors='coerce') # 2. 定义分箱区间和对应标签 bins = [0, 18, 30, 50, 65, 75, float('inf')] labels = ["18 and under", "19-30", "31-50", "51-65", "66-75", "75+"] # 3. 对数值型年龄分箱,保留缺失值 Students['AGE_BIN'] = pd.cut(Students['AGE'], bins=bins, labels=labels, include_lowest=True) # 4. 添加"Not_Available"类别,并填充缺失值 Students['AGE_BIN'] = Students['AGE_BIN'].cat.add_categories("Not_Available").fillna("Not_Available") # 5. 统计并输出各类别数量 age_counts = Students['AGE_BIN'].value_counts() for cat, num in age_counts.items(): print(f"{cat}: {num}")
代码说明
- 数值转换:用
pd.to_numeric配合errors='coerce',将无法转换的内容(含原NaN)统一转为NaN,避免类型冲突。 - 分箱逻辑:用
float('inf')覆盖75以上的所有年龄,确保无遗漏;include_lowest=True让0岁被归入第一个区间。 - 缺失值处理:先通过
cat.add_categories给分类列新增"Not_Available"类别,再用fillna将缺失值填充为该类别,保证分类完整性。
内容的提问来源于stack exchange,提问作者Mo_Phy
相关产品推荐
相关产品推荐

