You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas如何从存储字符串列表的列拆分生成对应Type的新列

Pandas按字符串类型拆分列表生成新列实现方案

首先修正你给出的示例代码的语法问题,原构造DataFrame的字典缺少外层大括号,正确的测试数据构造写法为:

import pandas as pd
df = pd.DataFrame({"A": [["Type1:Value1", "Type2:Value2", "Type1:Value3"]]})

方法1:直观易读版(适合小数据集)

直接逐行遍历拆分字符串,按类型归集值后生成新列,代码容错性高,就算Value内容本身包含冒号也不会拆分出错:

def parse_type_list(item_list):
    type_map = {}
    for item in item_list:
        # 只按第一个冒号拆分,避免Value中含冒号导致异常
        type_name, value = item.split(":", 1)
        type_map.setdefault(type_name, []).append(value)
    return pd.Series(type_map)

# 处理后合并回原DataFrame
df = pd.concat([df, df["A"].apply(parse_type_list)], axis=1)

# 不需要原A列可执行下行删除
# df.drop(columns=["A"], inplace=True)

方法2:高性能版(适合大数据集)

用pandas原生的explode、分组聚合操作实现,运行效率远高于apply循环,数据量大时优先选这个:

# 把每行的列表拆成独立行
exploded = df["A"].explode()
# 拆分出类型和值
split_df = exploded.str.split(":", n=1, expand=True).rename(columns={0:"Type", 1:"Value"})
# 按原行索引、类型分组,把同类型值聚合成列表后转成宽表
type_columns = split_df.groupby([split_df.index, "Type"])["Value"].apply(list).unstack()
# 合并回原表
df = pd.concat([df, type_columns], axis=1)

运行上述任意一种方法,都能得到你预期的结果:Type1列值为["Value1", "Value3"],Type2列值为["Value2"]。
如果你的5种Type是固定值,某行不存在对应Type时默认会填充NaN,需要替换为空列表的话,额外加一行处理即可:

# 替换成你实际的5个Type名称
fixed_types = ["Type1", "Type2", "Type3", "Type4", "Type5"]
df[fixed_types] = df[fixed_types].applymap(lambda x: x if isinstance(x, list) else [])

内容的提问来源于stack exchange,提问作者Jason

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 20:03:25