如何定义函数检查任意DataFrame的Age列并返回年龄分箱统计结果
错误原因
- 核心错误:
input()函数返回的是字符串类型,你输入的只是DataFrame变量的名称字符串,不是内存中实际的DataFrame对象,对字符串执行df['Age']的索引操作自然会抛出字符串索引必须为整数的TypeError。 - 语法错误:
pd.df['AgeGroup']写法错误,pd是pandas模块别名,不存在df属性,新增列应直接操作你自己的df变量,即df['AgeGroup']。 - 逻辑错误:代码中未定义
result变量,直接调用会抛出NameError;函数没有设计入参,通过输入变量名获取对象的方式属于反模式,完全不符合Python变量调用逻辑。 - 需求遗漏:没有实现各年龄分类的样本数量统计逻辑。
正确实现代码
import pandas as pd def age_range(df): # 入参合法性校验 if not isinstance(df, pd.DataFrame): raise TypeError("入参必须是pandas DataFrame类型") if 'Age' not in df.columns: raise ValueError("传入的DataFrame必须包含'Age'列") # 分箱规则增加inf边界,避免100岁以上数据出现空值 bins = [0, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, float('inf')] labels = ['0-9', '10-19', '20s', '30s', '40s', '50s', '60s', '70s', '80s', '90s', '100+'] # 年龄分箱打标 df['AgeGroup'] = pd.cut(df['Age'], bins=bins, labels=labels, right=False) # 统计各分组样本数量并按分组顺序排序 age_count = df['AgeGroup'].value_counts().sort_index() print("年龄分组样本统计:") print(age_count) # 支持返回分箱后的DataFrame和统计结果供后续使用 return df, age_count
使用示例
# 测试用例 test_df = pd.DataFrame({'Age': [5, 12, 25, 33, 47, 59, 62, 78, 85, 99, 102]}) processed_df, count_result = age_range(test_df)
内容的提问来源于stack exchange,提问作者Noob3000
相关产品推荐
相关产品推荐

