Pandas分组应用函数报错排查及超赞房东房价中位数差最优解法
Airbnb数据集:超赞/非超赞房东房价中位数差值分析
一、解决TypeError报错
你的报错是因为groupby.apply对多列分组时,会把整个分组的DataFrame作为单个参数传给函数,而不是拆分host_is_superhost和price为两个独立参数。
修复自定义函数
修改函数逻辑,让它接收分组后的DataFrame,再从中提取需要的列:
def func(group_df): # 筛选超赞/非超赞房东的房价数据 super_prices = group_df.loc[group_df['host_is_superhost'] == 't', 'price'] non_super_prices = group_df.loc[group_df['host_is_superhost'] == 'f', 'price'] # 处理分组中某类房东为空的情况,避免中位数计算报错 super_med = super_prices.median() if not super_prices.empty else 0 non_super_med = non_super_prices.median() if not non_super_prices.empty else 0 return super_med - non_super_med
正确调用代码
注意先把price列从带$的字符串转为数值类型:
listings = pd.read_csv("https://storage.googleapis.com/public-data-337819/listings%202%20reduced.csv", low_memory=False) # 清洗price列,转为浮点型 listings['price'] = listings['price'].replace('[\$,]', '', regex=True).astype(float) # 分组并计算差值 neighbourhood_groups = listings.groupby('neighbourhood_cleansed')[['host_is_superhost', 'price']] price_diffs = neighbourhood_groups.apply(func) # 找出差值最大的街区 top_neighbourhood = price_diffs.idxmax() top_diff_value = price_diffs.max()
二、更简洁的最优解法
无需自定义函数,用Pandas内置的pivot_table可以一步实现,代码更高效易读:
# 先清洗price列 listings['price'] = listings['price'].replace('[\$,]', '', regex=True).astype(float) # 生成透视表,按街区分组,计算两类房东的房价中位数 median_table = listings.pivot_table( index='neighbourhood_cleansed', columns='host_is_superhost', values='price', aggfunc='median' ) # 计算差值并筛选最大值 median_table['price_diff'] = median_table['t'] - median_table['f'] max_diff_result = (median_table['price_diff'].idxmax(), median_table['price_diff'].max()) print(f"差值最大的街区:{max_diff_result[0]},中位数差值:{max_diff_result[1]:.2f}")
该解法的优势
- 利用Pandas原生功能,避免自定义函数的冗余代码
- 透视表保留了两类房东的中位数数据,方便后续扩展分析
- 代码逻辑清晰,可读性和维护性更强
内容的提问来源于stack exchange,提问作者Advaita Mallik
相关产品推荐
相关产品推荐

