如何提取DataFrame中genres列的唯一电影类型列表?
提取DataFrame中|分隔的唯一电影类型
你的问题出在直接对genres列的完整字符串做去重,所以得到的是不同的类型组合,而非拆分后的单个类型。下面是两种可行的解决方法:
方法一:Pandas原生方法(推荐)
利用str.split拆分字符串,explode将列表展开为单行,再获取唯一值:
import pandas as pd # 拆分每个genres字符串为列表,再展开成单独行 single_genres = df['genres'].str.split('|').explode() # 获取唯一类型并转为列表 unique_genres = single_genres.unique().tolist() print(unique_genres)
方法二:结合字符串拼接与集合去重
将所有类型字符串拼接后拆分,再用集合或np.unique去重:
import numpy as np # 拼接所有genres字符串,按|拆分得到所有单个类型 all_genres = '|'.join(df['genres'].values).split('|') # 用集合去重(无序),或用np.unique排序(有序) unique_genres = list(set(all_genres)) # 有序版本 unique_genres_sorted = np.unique(all_genres).tolist() print(unique_genres_sorted)
执行上述代码后,就能得到所有单个的唯一电影类型,比如Action、Adventure、Science Fiction等。
内容的提问来源于stack exchange,提问作者possizola
相关产品推荐
相关产品推荐

