获取评分次数超50的餐厅名称时遇TypeError错误求助
问题:获取评分次数超过50次的餐厅列表并计算平均评分
原代码:
# Filter the rated restaurants df_rated = df[df['rating'] != 'Not given'].copy() # Convert rating column from object to integer df_rated['rating'] = df_rated['rating'].astype('int') # Create a dataframe that contains the restaurant names with their rating counts df_rating_count = df_rated.groupby(['restaurant_name'])['rating'].count().sort_values(ascending = False).reset_index() df_rating_count.head() # Get the restaurant names that have rating count more than 50 rest_names = df_rating_count[['rating']>50]['restaurant_name'] ## Complete the code to get the restaurant names having rating count more than 50 # Filter to get the data of restaurants that have rating count more than 50 df_mean_4 = df_rated[df_rated['restaurant_name'].isin(rest_names)].copy() # Group the restaurant names with their ratings and find the mean rating of each restaurant df_mean_4.groupby(['restaurant_name'])['rating'].mean().sort_values(ascending = False).reset_index().dropna() ## Complete the code to find the mean rating
运行报错:
TypeError Traceback (most recent call last) <ipython-input-46-9676daed1fbc> in <module> 1 # Get the restaurant names that have rating count more than 50 ----> 2 rest_names = df_rating_count[['rating']>50]['restaurant_name'] ## Complete the code to get the restaurant names having rating count more than 50 3 # Filter to get the data of restaurants that have rating count more than 50 4 df_mean_4 = df_rated[df_rated['restaurant_name'].isin(rest_names)].copy() 5 TypeError: '>' not supported between instances of 'list' and 'int'
错误原因
df_rating_count[['rating']]返回的是DataFrame类型(二维结构),无法直接与整数50做大小比较;而df_rating_count['rating']返回的是一维Series序列,可直接和整数做比较。
修正后的代码
# 过滤未评分餐厅 df_rated = df[df['rating'] != 'Not given'].copy() # 将评分列转为整数类型 df_rated['rating'] = df_rated['rating'].astype('int') # 统计每家餐厅的评分次数并降序排序 df_rating_count = df_rated.groupby(['restaurant_name'])['rating'].count().sort_values(ascending=False).reset_index() # 重命名计数列,提升可读性 df_rating_count.rename(columns={'rating': 'rating_count'}, inplace=True) # 获取评分次数超过50的餐厅名称 rest_names = df_rating_count[df_rating_count['rating_count'] > 50]['restaurant_name'] # 筛选目标餐厅的评分数据 df_mean_4 = df_rated[df_rated['restaurant_name'].isin(rest_names)].copy() # 计算平均评分并降序排序 result = df_mean_4.groupby(['restaurant_name'])['rating'].mean().sort_values(ascending=False).reset_index().dropna() # 重命名均值列 result.rename(columns={'rating': 'average_rating'}, inplace=True) print(result)
额外优化
可以合并步骤,减少中间变量,让逻辑更连贯高效:
# 一步完成:统计评分次数>50的餐厅,同时计算平均评分 result = df_rated.groupby('restaurant_name').agg( rating_count=('rating', 'count'), average_rating=('rating', 'mean') ).query('rating_count > 50').sort_values('average_rating', ascending=False).reset_index().dropna() print(result)
内容的提问来源于stack exchange,提问作者MauricioG
相关产品推荐
相关产品推荐

