如何将含16个变量的嵌套字典转换为Pandas DataFrame?
嵌套字典转Pandas DataFrame解决方案
问题背景
现有包含data(含多字段字典的列表)和summary(汇总统计字典)的嵌套字典,需转换为指定样式的Pandas DataFrame,原代码df = pd.DataFrame(dict.values())无法满足需求。
完整实现代码
import pandas as pd # 替换为你的原始字典变量 raw_dict = {'data': [{'direction_color': 'greenFont', 'rowDate': 'Apr 18, 2023', 'rowDateRaw': 1681776000, 'rowDateTimestamp': '2023-04-18T00:00:00Z', 'last_close': '1,268.550', 'last_open': '1,268.550', 'last_max': '1,268.550', 'last_min': '1,268.550', 'volume': '', 'volumeRaw': 0, 'change_precent': '0.32', 'last_closeRaw': '1268.55004882812500', 'last_openRaw': '1268.55004882812500', 'last_maxRaw': '1268.55004882812500', 'last_minRaw': '1268.55004882812500', 'change_precentRaw': 0.32187685232202745}, {'direction_color': 'greenFont', 'rowDate': 'Apr 17, 2023', 'rowDateRaw': 1681689600, 'rowDateTimestamp': '2023-04-17T00:00:00Z', 'last_close': '1,264.480', 'last_open': '1,264.480', 'last_max': '1,264.480', 'last_min': '1,264.480', 'volume': '', 'volumeRaw': 0, 'change_precent': '1.03', 'last_closeRaw': '1264.47998046875000', 'last_openRaw': '1264.47998046875000', 'last_maxRaw': '1264.47998046875000', 'last_minRaw': '1264.47998046875000', 'change_precentRaw': 1.0274672346024991}], 'summary': {'last_highest': '1,268.550', 'last_lowest': '1,264.480', 'last_difference': '4.070', 'last_avarage_total': '1,266.515', 'last_change_percent': '1.353'}} # 1. 从data列表生成基础DataFrame df = pd.DataFrame(raw_dict['data']) # 2. 筛选目标列并设置中文列名 df = df[['rowDate', 'last_open', 'last_max', 'last_min', 'last_close', 'change_precent']] df.columns = ['日期', '开盘价', '最高价', '最低价', '收盘价', '涨跌幅(%)'] # 3. 清理数值格式,转换为数值类型 numeric_cols = ['开盘价', '最高价', '最低价', '收盘价'] df[numeric_cols] = df[numeric_cols].apply(lambda col: col.str.replace(',', '').astype(float)) df['涨跌幅(%)'] = df['涨跌幅(%)'].astype(float) # 4. 创建汇总行并合并到主表 summary_row = pd.DataFrame([{ '日期': '汇总', '开盘价': '-', '最高价': float(raw_dict['summary']['last_highest'].replace(',', '')), '最低价': float(raw_dict['summary']['last_lowest'].replace(',', '')), '收盘价': '-', '涨跌幅(%)': float(raw_dict['summary']['last_change_percent']) }]) final_df = pd.concat([df, summary_row], ignore_index=True) # 输出结果 print(final_df)
代码说明
- 提取核心数据:直接用字典的
data列表生成DataFrame,这是目标表格的主体数据。 - 列筛选与重命名:保留需要的字段,设置直观的中文列名匹配需求样式。
- 数值格式处理:移除数值中的千分位逗号,转换为浮点型,确保数据可正常计算或展示。
- 添加汇总行:从
summary字段提取汇总数据,构建单独的DataFrame后合并到主表底部,完成最终样式。
原代码问题分析
pd.DataFrame(dict.values())会将data和summary作为两行数据,把每个子字段作为列,完全不符合目标表格的结构,因此需要拆分处理核心数据与汇总数据。
内容的提问来源于stack exchange,提问作者Sebby
相关产品推荐
相关产品推荐

