如何基于DataFrame列与Series值创建条件新DataFrame
实现方法
方法一:简单直观的apply方法
先构造示例测试数据:
import pandas as pd # 构造输入的Series s = pd.Series( [['A', 'B', 'C'], ['D', 'E'], ['B', 'C', 'D']], index=['t1', 't2', 't3'] ) # 目标列名列表 target_cols = ['A', 'B', 'C', 'D', 'E']
通过apply遍历Series的每个列表元素,对每个目标列名判断是否存在,再转为Series:
# 生成布尔值结果 result_df = s.apply(lambda lst: pd.Series([col in lst for col in target_cols], index=target_cols)) # 如果需要显示'T'/'F'而非布尔值,替换为: # result_df = s.apply(lambda lst: pd.Series(['T' if col in lst else 'F' for col in target_cols], index=target_cols))
运行后得到的result_df结构如下:
A B C D E t1 True True True False False t2 False False False True True t3 False True True True False
方法二:高效向量化方法(适合大数据集)
如果数据量较大,apply的循环效率较低,可以用explode+pivot的向量化方案:
# 将Series的列表元素展开为单行一个列名 exploded = s.explode().reset_index(name='column') # 标记该列名存在 exploded['exists'] = True # 透视成目标结构,缺失的列填充为False result_df = exploded.pivot(index='index', columns='column', values='exists') # 对齐目标列,填充缺失值 result_df = result_df.reindex(columns=target_cols, fill_value=False) # 同样,如需转为'T'/'F': # result_df = result_df.replace({True: 'T', False: 'F'})
两种方法都能得到你需要的DataFrame,根据数据规模选择即可。
内容的提问来源于stack exchange,提问作者WWW
相关产品推荐
相关产品推荐

