基于DataFrame分组在Python中构建Dijkstra模型所需距离矩阵
基于分组生成Dijkstra所需的距离矩阵
需求说明
我要搭建Dijkstra模型,需要基于以下DataFrame(附创建代码),按照column1的取值与person字段进行分组,为每组生成对应的距离矩阵。例如当column1为1且person为A时,需生成[[0,20],[40,0]]形式的矩阵。
原始数据与DataFrame创建
import pandas as pd data={'column1':[1,1,1,1,1,1,1,1,2,2,2,2],'person':['A','A','A','A','B','B','B','B','C','C','C','C'],'location1':['GOA','BANGLORE','GOA','BANGLORE','BANGLORE','DELHI','BANGLORE','DELHII','KOCHI','DELHI','DELHI','KOCHI'],'location2':['BANGLORE','GOA','GOA','BANGLORE','DELHI','DELHI','BANGLORE','BANGLORE','DELHI','KOCHI','DELHI','KOCHI'],'time':[20,40,0,0,34,0,0,23,21,56,0,0]} df = pd.DataFrame(data)
解决方案步骤
1. 修正数据错误
原始数据中person为B的组里存在拼写错误DELHII,先修正为DELHI,避免地点分类混乱:
df['location1'] = df['location1'].replace('DELHII', 'DELHI')
2. 定义生成距离矩阵的函数
该函数会针对每个分组提取唯一地点,构建对称的距离矩阵:
def create_distance_matrix(group): # 获取分组内所有唯一地点并排序,保证矩阵顺序固定 locations = sorted(group[['location1', 'location2']].stack().unique()) loc_index = {loc: idx for idx, loc in enumerate(locations)} n = len(locations) # 初始化全0矩阵 matrix = [[0]*n for _ in range(n)] # 填充矩阵对应位置的时间值 for _, row in group.iterrows(): i = loc_index[row['location1']] j = loc_index[row['location2']] matrix[i][j] = row['time'] return matrix
3. 分组生成矩阵
按column1和person分组,应用上述函数生成各组的距离矩阵:
# 结果以字典存储,键为(column1, person)元组,值为对应距离矩阵 distance_matrices = df.groupby(['column1', 'person']).apply(create_distance_matrix).to_dict()
4. 查看结果
# 查看column1=1、person=A的矩阵 print(distance_matrices[(1, 'A')]) # 输出:[[0, 20], [40, 0]] # 查看column1=1、person=B的矩阵 print(distance_matrices[(1, 'B')]) # 输出:[[0, 34], [23, 0]] # 查看column1=2、person=C的矩阵 print(distance_matrices[(2, 'C')]) # 输出:[[0, 21], [56, 0]]
补充说明
- 对地点进行排序是为了确保每次生成的矩阵顺序一致,避免因原始数据顺序导致矩阵结构混乱
- 矩阵对角线默认保持为0,对应地点到自身的距离
内容的提问来源于stack exchange,提问作者Mathew John
相关产品推荐
相关产品推荐

