You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pandas将含列表、字典的dataframe列转换为独立展开列的方法

解决方案

核心思路

优先保证代码可读性和兼容性,先统一解析不同格式的字符串值为Python原生对象,再按行映射生成展开后的结构,最后合并为结果表。

实现代码

首先导入依赖库:

import pandas as pd
import ast
import numpy as np

编写行处理函数完成单条数据的解析和展开:

def expand_row(row):
    # 统一解析shop_id为列表格式
    try:
        shop_list = ast.literal_eval(row["shop_id"])
        shop_list = shop_list if isinstance(shop_list, list) else [shop_list]
    except:
        # 兼容单个字符串形式的shop_id
        shop_list = [row["shop_id"]]
    
    # 解析price生成shop_id到价格的映射
    price_map = {}
    try:
        price_raw = ast.literal_eval(row["price"])
        if isinstance(price_raw, dict):
            # 字典格式反转键值生成shop->price映射
            for price, shops in price_raw.items():
                for s in shops:
                    price_map[s] = price
        elif pd.isna(price_raw) or str(price_raw) == "NaN":
            # 空值场景
            for s in shop_list:
                price_map[s] = np.nan
        else:
            # 单个价格场景
            for s in shop_list:
                price_map[s] = str(price_raw)
    except:
        # 兼容纯数字字符串形式的价格
        for s in shop_list:
            price_map[s] = row["price"]
    
    # 返回当前行对应的展开后子表
    return pd.DataFrame({
        "item_id": row["item_id"],
        "shop_id": shop_list,
        "price": [price_map[s] for s in shop_list]
    })

调用执行得到结果:

result = pd.concat(df.apply(expand_row, axis=1).tolist(), ignore_index=True)

方案说明

  • 用ast.literal_eval解析字符串化的列表/字典,比eval更安全,只会解析Python字面量,不会执行恶意代码,是这类场景的标准处理方案。
  • 内置异常捕获兼容非结构化的单个取值,不需要提前做全列格式校验,适配性更强。
  • 行处理逻辑可维护性高,后续如果新增取值格式,只需要修改expand_row函数即可。
  • 十万级以下数据量用该方案性能完全够用,如果是百万级以上的大规模数据集,可以把解析逻辑改为pandas的向量化str操作,减少apply的开销。

内容的提问来源于stack exchange,提问作者charlie_boy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 00:39:04