Python DataFrame转Dict:适配配送取货双节点模型的优化需求
问题描述
我有个记录节点间通行成本的DataFrame:
Base_Node Turbine_1 Turbine_2 Charging_Station_1 0 Base_Node 0.0 1.1 5.3 23.5 1 Turbine_1 1.1 0.0 4.2 23.4 2 Turbine_2 5.3 4.2 0.0 22.8 3 Charging_Station_1 23.5 23.4 22.8 0.0
这个矩阵里的数值代表从节点i到节点j的通行成本。之前我写了个把它转成字典的函数,但现在要适配配送/取货场景:每个节点得拆成配送、取货两个节点,索引规则是:
- Base Node:配送起点对应索引
0,取货终点对应索引2n+1(n是原非Base节点的数量,这里n=3) - Turbine、Charging Station这类非Base节点:配送节点占索引
1~n,取货节点占n+1~2n,同一个原节点的配送节点i和取货节点对应i和i+n
原转换函数是这样的:
def matrix_to_dict(cost_df): cost_of_route = {} for row in range(cost_df.shape[0]): # 遍历行 for col in range(1, cost_df.shape[1]): # 跳过第一列的节点名 cost_of_route[(row, col-1)] = cost_df.iloc[row, col] return cost_of_route
这函数满足不了双节点的映射需求,我自己写了新代码实现了功能,但效率极低,现在求更高效的实现方案。我的低效代码如下:
def matrix_to_dict(cost_df,n): cost_of_route = {} for row in range(0, 1): # Loop through rows for col in range(1, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(row, col-1)] = cost_df.iloc[row, col] for row in range(0, 1): # Loop through rows for col in range(1, cost_df.shape[1]-1): # Start from 1 to exclude first column cost_of_route[(row, col+n)] = cost_df.iloc[row, col+1] for row in range(0, cost_df.shape[0]): # Loop through rows for col in range(1, 2): # Start from 1 to exclude first column cost_of_route[(row, col-1)] = cost_df.iloc[row, col] for row in range(0, cost_df.shape[0]): # Loop through rows for col in range(1, 2): # Start from 1 to exclude first column cost_of_route[(row+n, col-1)] = cost_df.iloc[row, col] for row in range(1, cost_df.shape[0]): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(row, col-1)] = cost_df.iloc[row, col] for row in range(1, cost_df.shape[0]): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(row+n, col-1)] = cost_df.iloc[row, col] for row in range(1, cost_df.shape[0]): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(row, col-1+n)] = cost_df.iloc[row, col] for row in range(1, cost_df.shape[0]): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 2 to exclude first column cost_of_route[(row+n, col-1+n)] = cost_df.iloc[row, col] for row in range(0, 1): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(n+n+1, col-1+n)] = cost_df.iloc[row, col] # print(n+n+1, col-1+n) for row in range(0, 1): # Loop through rows for col in range(2, cost_df.shape[1]): # Start from 1 to exclude first column cost_of_route[(col-1+n, n+n+1)] = cost_df.iloc[row, col] # print(col-1+n, n+n+1) for row in range(0, cost_df.shape[0]): # Loop through rows for col in range(1, 2): # Start from 1 to exclude first column cost_of_route[(row+n, col-1+n)] = cost_df.iloc[row, col] # print(row+n, col-1+n) for row in range(0, cost_df.shape[0]): # Loop through rows for col in range(1, 2): # Start from 1 to exclude first column cost_of_route[(col-1+n, row+n)] = cost_df.iloc[row, col] return cost_of_route
预期输出格式示例(对应n=3的情况):
{(0, 0): '0.0', (0, 1): '21.4', (0, 2): '26.4', (0, 3): '21.6', (0, 4): '21.4', (0, 5): '26.4', (0, 6): '21.6', (1, 0): '21.4', (2, 0): '26.4', (3, 0): '0.0', (4, 0): '21.4', (5, 0): '26.4', (6, 0): '21.6', (1, 1): '0.0', (1, 2): '5.0', (1, 3): '1.3', (2, 1): '5.0', (2, 2): '0.0', (2, 3): '5.0', (3, 1): '1.3', (3, 2): '5.0', (3, 3): '0.0', (4, 1): '0.0', (4, 2): '5.0', (4, 3): '21.4', (5, 1): '5.0', (5, 2): '0.0', (5, 3): '26.4', (6, 1): '1.3', (6, 2): '5.0', (6, 3): '21.6', (1, 4): '0.0', (1, 5): '5.0', (1, 6): '1.3', (2, 4): '5.0', (2, 5): '0.0', (2, 6): '5.0', (3, 4): '21.4', (3, 5): '26.4', (3, 6): '21.6', (4, 4): '0.0', (4, 5): '5.0', (4, 6): '1.3', (5, 4): '5.0', (5, 5): '0.0', (5, 6): '5.0', (6, 4): '1.3', (6, 5): '5.0', (6, 6): '0.0', (7, 4): '21.4', (7, 5): '26.4', (7, 6): '21.6', (4, 7): '21.4', (5, 7): '26.4', (6, 7): '21.6'}
高效实现方案
核心思路就是先把原节点和新节点的映射关系理清楚,然后只遍历一次原矩阵,把所有需要的新节点成本对都填到字典里,避免原代码里那种重复循环的冗余操作。
直接上高效代码:
def matrix_to_dict(cost_df): cost_dict = {} # 原节点总数,n是原非Base节点的数量 total_nodes = cost_df.shape[0] n = total_nodes - 1 end_node = 2 * n + 1 # Base节点的取货终点索引 # 提取纯成本数值矩阵,避免多次调用iloc拖慢速度 cost_matrix = cost_df.iloc[:, 1:].values # 遍历所有原节点对(i,j),一次循环搞定所有映射 for i in range(total_nodes): for j in range(total_nodes): cost = cost_matrix[i, j] # 处理原节点i对应的新节点 if i == 0: # 原Base节点对应新配送起点0 i_delivery = 0 else: # 非Base节点的配送、取货节点索引 i_delivery = i i_pickup = i + n # 处理原节点j对应的新节点 if j == 0: # 原Base节点对应新取货终点end_node j_pickup = end_node else: # 非Base节点的配送、取货节点索引 j_delivery = j j_pickup = j + n # 场景1:从i的配送节点到j的配送节点 if i == 0: if j != 0: cost_dict[(i_delivery, j_delivery)] = cost else: cost_dict[(i_delivery, j_delivery)] = cost # 场景2:从i的配送节点到j的取货节点(j不是Base节点) if j != 0: if i == 0: cost_dict[(i_delivery, j_pickup)] = cost else: cost_dict[(i_delivery, j_pickup)] = cost # 场景3:从i的取货节点到j的配送节点(i不是Base节点) if i != 0: if j != 0: cost_dict[(i_pickup, j_delivery)] = cost else: cost_dict[(i_pickup, j_pickup)] = cost_matrix[j, i] # 对应取货节点到Base终点的成本 # 场景4:从i的取货节点到j的取货节点(i、j都不是Base节点) if i != 0 and j != 0: cost_dict[(i_pickup, j_pickup)] = cost # 补充Base终点到取货节点的成本,以及取货节点到Base终点的成本 if i == 0 and j != 0: # Base终点到j取货节点的成本 = 原j到Base的成本 cost_dict[(end_node, j_pickup)] = cost_matrix[j, 0] # j取货节点到Base终点的成本 = 原Base到j的成本 cost_dict[(j_pickup, end_node)] = cost # 补充Base节点自身的映射 cost_dict[(0, 0)] = cost_matrix[0, 0] cost_dict[(end_node, end_node)] = cost_matrix[0, 0] return cost_dict
关键点说明
- 直接用numpy数组:把DataFrame里的纯数值提取成numpy矩阵,比反复调用
pandas.iloc快得多 - 一次遍历全覆盖:只用一次双重循环遍历原矩阵的所有(i,j)对,一次性生成配送→配送、配送→取货、取货→配送、取货→取货这四种场景的成本映射,同时处理Base节点的特殊规则
- 避免冗余逻辑:把原节点到新节点的映射规则集中处理,不像原代码那样拆成多个循环,减少重复计算
- 适配预期输出:补充了Base终点和取货节点之间的双向成本映射,完全匹配你给出的预期输出格式
用你提供的DataFrame测试这个函数,结果会和预期一致,而且节点数量越多,效率提升越明显。
内容的提问来源于stack exchange,提问作者Rune tønnessen
相关产品推荐
相关产品推荐

