如何使用pandas从电影DataFrame中统计美国排名前5的电影类型
解决方案
你原来的代码有两个核心问题:一是groupby(['country']['listed_in'])属于语法错误,不符合pandas多列分组的写法;二是listed_in存储的是逗号分隔的多类型字符串,直接分组统计只会把整段字符串当一个类型,结果不符合需求。可按以下步骤实现:
1. 筛选美国地区影片
先过滤出国家包含美国的条目,可根据需求选择匹配规则:
# 规则1:包含美国的合拍影片也统计在内 us_movies = netflix_df[netflix_df['country'].fillna('').str.contains('United States')].copy() # 规则2:仅统计单独标注为美国的影片,替换上一行即可 # us_movies = netflix_df[netflix_df['country'] == 'United States'].copy()
2. 拆分类型字段并展开
将多值的类型字符串拆分为列表,再用explode方法转为一行对应一个类型的结构:
# 按「逗号+空格」拆分类型为列表 us_movies['listed_in'] = us_movies['listed_in'].str.split(', ') # 展开类型列,每个类型单独占一行 us_type_expand = us_movies.explode('listed_in')
3. 统计排名前5的类型
top5_us_types = us_type_expand['listed_in'].value_counts().sort_values(ascending=False).head(5) print(top5_us_types)
也可以用链式写法一行完成:
top5_us_types = netflix_df[netflix_df['country'].fillna('').str.contains('United States')]\ .assign(listed_in=lambda x: x['listed_in'].str.split(', '))\ .explode('listed_in')['listed_in']\ .value_counts()\ .head(5)
内容的提问来源于stack exchange,提问作者Juliette
相关产品推荐
相关产品推荐

