You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将字典列表转换为保持键相对顺序的Pandas DataFrame

问题:将有序字典列表转换为Pandas DataFrame时保留键的相对顺序

给定有序字典列表(Python 3.7+ 字典默认保留插入顺序),转换为Pandas DataFrame时,默认生成的列顺序无法满足需求:需要保证所有字典中键的相对顺序一致——即如果在任意一个字典中键X出现在Y之前,那么最终DataFrame的列中X必须在Y之前;无直接顺序约束的键,按首次出现顺序排列。

示例输入:

from typing import List, Dict, Any

testDict: List[Dict[str, Any]] = [
    {"A": 0.1, "B": 1, "E": "ABE"},
    {"A": 0.11, "B": 20, "C": 0.2, "E": "ABCE"},
    {"A": 0.11, "B": 3, "D": 33, "E": "ABDE"},
    {"A": 0.13, "B": 40, "C": 0.5, "D": 23, "E": "ABCDE"},
]

使用pd.json_normalize转换后列顺序为['A','B','E','C','D'],但期望列顺序为['A','B','C','D','E'];若移除最后一行,期望列顺序为['A','B','D','C','E'](按无约束键的首次出现顺序)。

另一示例输入:

testDict2 = [
    {'XYZ': 0.1, 'ABC': 1, 'PQR': 'ABE'},
    {'XYZ': 0.11, 'ABC': 20, 'KLM': 0.2, 'PQR': 'ABCE'},
    {'XYZ': 0.11, 'ABC': 3, 'DEF': 33, 'PQR': 'ABDE'},
    {'XYZ': 0.13, 'ABC': 40, 'KLM': 0.5, 'DEF': 23, 'PQR': 'ABCDE'},
]

期望列顺序为['XYZ','ABC','KLM','DEF','PQR']。

实际场景为10万行数据,共256个键,需要高效实现需求。

解决方案

通过拓扑排序构建符合所有顺序约束的列顺序,同时利用键的首次出现顺序解决无约束键的排序问题,具体实现如下:

import pandas as pd
from collections import defaultdict, deque
from typing import List, Dict, Any

def get_column_order(dicts_list: List[Dict[str, Any]]) -> List[str]:
    # 构建邻接表(记录键的顺序约束)和入度字典
    adjacency = defaultdict(set)
    in_degree = defaultdict(int)
    # 记录所有键、键的首次出现顺序(用于无约束时的排序优先级)
    all_keys = set()
    first_occur_idx = {}
    current_idx = 0

    for d in dicts_list:
        keys = list(d.keys())
        # 更新首次出现顺序
        for key in keys:
            if key not in first_occur_idx:
                first_occur_idx[key] = current_idx
                current_idx += 1
            all_keys.add(key)
        # 添加相邻键的顺序约束(X必须在Y之前)
        for i in range(len(keys) - 1):
            x, y = keys[i], keys[i+1]
            if y not in adjacency[x]:
                adjacency[x].add(y)
                in_degree[y] += 1

    # 初始化拓扑排序队列:入度为0的键,按首次出现顺序排序
    queue = deque(sorted(
        [k for k in all_keys if in_degree.get(k, 0) == 0],
        key=lambda x: first_occur_idx[x]
    ))
    column_order = []

    while queue:
        node = queue.popleft()
        column_order.append(node)
        # 更新邻接键的入度,入度为0则加入队列
        for neighbor in adjacency[node]:
            in_degree[neighbor] -= 1
            if in_degree[neighbor] == 0:
                queue.append(neighbor)

    return column_order

# 测试第一个示例
testDict = [
    {"A": 0.1, "B": 1, "E": "ABE"},
    {"A": 0.11, "B": 20, "C": 0.2, "E": "ABCE"},
    {"A": 0.11, "B": 3, "D": 33, "E": "ABDE"},
    {"A": 0.13, "B": 40, "C": 0.5, "D": 23, "E": "ABCDE"},
]
cols_order = get_column_order(testDict)
testDf = pd.DataFrame(testDict)[cols_order]
print(testDf)
# 输出列顺序:A, B, C, D, E

# 测试第二个示例
testDict2 = [
    {'XYZ': 0.1, 'ABC': 1, 'PQR': 'ABE'},
    {'XYZ': 0.11, 'ABC': 20, 'KLM': 0.2, 'PQR': 'ABCE'},
    {'XYZ': 0.11, 'ABC': 3, 'DEF': 33, 'PQR': 'ABDE'},
    {'XYZ': 0.13, 'ABC': 40, 'KLM': 0.5, 'DEF': 23, 'PQR': 'ABCDE'},
]
cols_order2 = get_column_order(testDict2)
testDf2 = pd.DataFrame(testDict2)[cols_order2]
print(testDf2)
# 输出列顺序:XYZ, ABC, KLM, DEF, PQR

性能说明

  • 遍历10万行数据时,仅需处理键的顺序约束,集合操作的开销极低;
  • 拓扑排序仅针对256个键,计算量可忽略不计;
  • 最终转换DataFrame时,直接按预先生成的列顺序索引,效率与常规DataFrame构造一致。

内容的提问来源于stack exchange,提问作者soumeng78

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 16:10:26