使用pandas将含列表、字典的dataframe列转换为独立展开列的方法
解决方案
核心思路
优先保证代码可读性和兼容性,先统一解析不同格式的字符串值为Python原生对象,再按行映射生成展开后的结构,最后合并为结果表。
实现代码
首先导入依赖库:
import pandas as pd import ast import numpy as np
编写行处理函数完成单条数据的解析和展开:
def expand_row(row): # 统一解析shop_id为列表格式 try: shop_list = ast.literal_eval(row["shop_id"]) shop_list = shop_list if isinstance(shop_list, list) else [shop_list] except: # 兼容单个字符串形式的shop_id shop_list = [row["shop_id"]] # 解析price生成shop_id到价格的映射 price_map = {} try: price_raw = ast.literal_eval(row["price"]) if isinstance(price_raw, dict): # 字典格式反转键值生成shop->price映射 for price, shops in price_raw.items(): for s in shops: price_map[s] = price elif pd.isna(price_raw) or str(price_raw) == "NaN": # 空值场景 for s in shop_list: price_map[s] = np.nan else: # 单个价格场景 for s in shop_list: price_map[s] = str(price_raw) except: # 兼容纯数字字符串形式的价格 for s in shop_list: price_map[s] = row["price"] # 返回当前行对应的展开后子表 return pd.DataFrame({ "item_id": row["item_id"], "shop_id": shop_list, "price": [price_map[s] for s in shop_list] })
调用执行得到结果:
result = pd.concat(df.apply(expand_row, axis=1).tolist(), ignore_index=True)
方案说明
- 用
ast.literal_eval解析字符串化的列表/字典,比eval更安全,只会解析Python字面量,不会执行恶意代码,是这类场景的标准处理方案。 - 内置异常捕获兼容非结构化的单个取值,不需要提前做全列格式校验,适配性更强。
- 行处理逻辑可维护性高,后续如果新增取值格式,只需要修改
expand_row函数即可。 - 十万级以下数据量用该方案性能完全够用,如果是百万级以上的大规模数据集,可以把解析逻辑改为pandas的向量化str操作,减少apply的开销。
内容的提问来源于stack exchange,提问作者charlie_boy
相关产品推荐
相关产品推荐

