如何展开Pandas数据表嵌套列表列并转为独立列
Pandas嵌套列表列转布尔独立列解决方案
示例数据构造
先模拟你的数据结构,方便测试验证:
import pandas as pd data = { 'id': [1, 2, 3], 'mail': ['a@test.com', 'b@test.com', 'c@test.com'], 'displayName': ['Alice', 'Bob', 'Charlie'], 'propertiesRegistered': [ ['address', 'mobilePhone'], ['officePhone'], ['mobilePhone', 'officePhone'] ], 'createdDateTime': ['2024-01-01', '2024-01-02', '2024-01-03'] } df = pd.DataFrame(data)
方法一:快速生成布尔列(推荐)
利用str.get_dummies一键生成哑变量后转布尔值,再和原表合并:
# 将列表转为逗号分隔字符串,生成哑变量并转布尔类型 dummy_cols = df['propertiesRegistered'].str.join(',').str.get_dummies(sep=',').astype(bool) # 合并原数据(移除原嵌套列表列)和新生成的布尔列 final_df = pd.concat([df.drop('propertiesRegistered', axis=1), dummy_cols], axis=1)
方法二:自定义属性列(灵活可控)
如果需要明确指定要生成的属性列(或提前知道所有可能属性),用遍历方式生成:
# 提取所有唯一属性值(若提前知道属性列表,直接写死会更高效) all_props = set(prop for sublist in df['propertiesRegistered'] for prop in sublist) # 为每个属性生成对应的布尔列 for prop in all_props: df[prop] = df['propertiesRegistered'].apply(lambda x: prop in x) # 移除原嵌套列表列,得到最终结果 final_df = df.drop('propertiesRegistered', axis=1)
最终输出示例
id mail displayName createdDateTime address mobilePhone officePhone 0 1 a@test.com Alice 2024-01-01 True True False 1 2 b@test.com Bob 2024-01-02 False False True 2 3 c@test.com Charlie 2024-01-03 False True True
内容的提问来源于stack exchange,提问作者KrunkFu
相关产品推荐
相关产品推荐

