如何查找单个DataFrame中top30与played_games两列的共同元素
问题解决说明
你提供的原始数据框、错误输出、期望输出如下:
你当前得到逐字符拆分的错误结果,是因为played_games和top 30两列中看起来是列表的内容,实际存储为字符串类型。直接对字符串执行set()操作会将其拆分为单个独立字符,最终得到不符合预期的交集结果。
解决步骤
- 首先导入
ast模块,用安全的方法将字符串格式的列表转换为真实列表对象:
import ast # 对两列完成格式转换 Prediction1['played_games'] = Prediction1['played_games'].apply(ast.literal_eval) Prediction1['top 30'] = Prediction1['top 30'].apply(ast.literal_eval)
- 再执行交集计算逻辑即可得到正确结果,如果需要输出为列表而非集合,可额外套一层
list():
# 输出为集合格式 Prediction1['precision_at_30'] = [ set(a).intersection(b) for a, b in zip(Prediction1['played_games'], Prediction1['top 30']) ] # 如果需要输出为列表格式,使用以下代码 # Prediction1['precision_at_30'] = [ # list(set(a).intersection(b)) for a, b in zip(Prediction1['played_games'], Prediction1['top 30']) # ]
内容的提问来源于stack exchange,提问作者lifo
相关产品推荐
相关产品推荐

