You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何高效基于滞后坐标列排序路线点数据集并生成序列

高效生成路线点序列的方案(替代逐行遍历)

针对你遇到的问题——按Day Type-Direction分组,通过(to)与(from)坐标匹配还原路线点顺序,完全不用逐行遍历,这里有两种高效实现思路:

方案1:用NetworkX构建有向图找路径

路线本质是一条无分支的线性有向链,用图论工具可以快速定位顺序:

  • 把每个点的(from)坐标作为节点,(to)坐标作为目标节点,给每个组构建有向边
  • 找到入度为0的节点(就是路线起点,没有任何点的(to)坐标能匹配它的(from))
  • 从起点出发提取完整路径,按路径顺序排序组内数据即可

示例代码:

import pandas as pd
import networkx as nx

# 给每个点生成唯一标识(用from坐标组合)
df['from_node'] = df.apply(lambda x: (x['Latitude (from)'], x['Longitude (from)']), axis=1)
df['to_node'] = df.apply(lambda x: (x['Latitude (to)'], x['Longitude (to)']), axis=1)

final_list = []
for group_key, sub_df in df.groupby(['Day Type', 'Direction']):
    # 构建有向图
    graph = nx.DiGraph()
    graph.add_edges_from(zip(sub_df['from_node'], sub_df['to_node']))
    # 定位起点:入度为0的节点(路线的第一个点)
    start = [node for node, degree in graph.in_degree() if degree == 0][0]
    # 获取完整路径(单链场景下路径唯一)
    path = nx.shortest_path(graph, start)
    # 按路径顺序排序并添加序列列
    sorted_sub = sub_df.set_index('from_node').loc[path].reset_index()
    sorted_sub['sequence'] = range(1, len(sorted_sub)+1)
    final_list.append(sorted_sub)

# 合并所有组的结果
final_df = pd.concat(final_list, ignore_index=True)

方案2:纯Pandas映射实现(无需额外依赖)

如果不想安装第三方库,用Pandas自身的映射关系就能搞定:

  • 先给每个组构建to_node到from_node的映射表
  • 找到起点(from_node不在所有to_node集合里的点)
  • 从起点开始循环查找下一个节点,生成完整路径后排序

示例代码:

import pandas as pd

# 生成节点标识
df['from_node'] = df.apply(lambda x: (x['Latitude (from)'], x['Longitude (from)']), axis=1)
df['to_node'] = df.apply(lambda x: (x['Latitude (to)'], x['Longitude (to)']), axis=1)

final_list = []
for group_key, sub_df in df.groupby(['Day Type', 'Direction']):
    # 构建to节点到from节点的映射
    to_to_from = sub_df.set_index('to_node')['from_node'].to_dict()
    # 找起点:from节点不在to节点列表里的那个点
    start = sub_df[~sub_df['from_node'].isin(sub_df['to_node'])]['from_node'].iloc[0]
    # 生成路径
    path = [start]
    current_node = start
    while current_node in to_to_from:
        current_node = to_to_from[current_node]
        path.append(current_node)
    # 排序并添加序列列
    sorted_sub = sub_df.set_index('from_node').loc[path].reset_index()
    sorted_sub['sequence'] = range(1, len(sorted_sub)+1)
    final_list.append(sorted_sub)

final_df = pd.concat(final_list, ignore_index=True)

这两种方法的效率都远高于逐行遍历,数据量越大优势越明显:NetworkX的方案适合扩展到更复杂的路线场景(比如偶尔有分支),纯Pandas方案则更轻量,不需要额外安装库。

内容的提问来源于stack exchange,提问作者z star

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 07:25:29