You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升DataFrame中列表元素的搜索速度?

优化地理坐标匹配的向量化方案

我有一个按输电线路分组、包含地理坐标列表列的DataFrame,还有一个独立的坐标列表。当样本编号为DS_3时,需要找到每个坐标在DataFrame中对应的输电线路ID,但近6000个坐标用循环遍历的方式耗时极长。我知道循环DataFrame效率很低,推测有向量化实现方法,但对这类方法不太熟悉,求优化建议。


数据示例

DataFrame结构

id  coordinates                                 voltage   length    voltage x length
0   36  [[-81.443569, 28.470022], [-81.446726, 28.4740...   230 2788.450481 6.413436e+05
1   69  [[-82.402208, 27.907588], [-82.406592, 27.9084...   69  1486.634968 1.025778e+05
2   87  [[-80.38392, 25.748665], [-80.383758, 25.74358...   69  3395.795388 2.343099e+05
3   128 [[-81.423956, 28.410278], [-81.424811, 28.4053...   69  5231.189711 3.609521e+05
4   138 [[-81.843314, 30.572359], [-81.844404, 30.5685...   230 2716.984353 6.249064e+05
... ... ... ... ... ...
3061    68184   [[-81.251491, 28.53718], [-81.250396, 28.53283...   69  19243.512450    1.327802e+06
3062    68189   [[-82.669886, 28.961533], [-82.664782, 28.9615...   230 27463.901761    6.316697e+06
3063    68196   [[-81.157196, 29.000982], [-81.157041, 28.9958...   500 90524.038042    4.526202e+07
3064    68199   [[-80.549594, 28.481094], [-80.551733, 28.4857...   115 7185.881445 8.263764e+05
3065    68211   [[-80.44025, 25.81403]] 115 673.881802  7.749641e+04

独立坐标列表示例

[[-81.708274, 31.095992], [-81.708763, 31.090911], [-81.709349, 31.085841], [-81.710002, 31.080779], [-81.710627, 31.075713], [-81.711167, 31.070638], [-81.711649, 31.065557], [-81.712316, 31.060497], [-81.713036, 31.055444], [-81.713757, 31.050391], [-81.714478, 31.045338], [-81.715199, 31.040285], [-82.184384, 31.058297], [-82.188367, 31.061488], [-82.192045, 31.065027], [-82.195735, 31.068554], [-82.199426, 31.072079], [-82.20315, 31.075567], [-82.207127, 31.078767], [-82.211101, 31.081969], [-82.215077, 31.08517], [-82.219057, 31.088366], [-82.223033, 31.091567], [-82.227002, 31.094776], [-82.230978, 31.097977], [-82.234959, 31.101171], [-82.238934, 31.104373], [-82.242912, 31.107571], [-82.24689, 31.110769], [-82.250862, 31.113975], [-82.255483, 31.116065], [-82.260227, 31.117947], [-82.26497, 31.119832], [-82.269711, 31.121722], [-82.274457, 31.123602], [-82.279199, 31.125489], [-82.283947, 31.127364], [-82.28869, 31.129249], [-82.293435, 31.131131], [-82.298182, 31.133006], [-82.30292, 31.134905] ...  

现有低效代码

初始版本

for index, row in df.iterrows():
    for i in range(len(coordinates)):
        if coordinates[i] in row['coordinates'] and sample[i]['sample_0'] == 'DS_3':
                    if row['id'] in damaged_lines:
                        break
                    else:
                        damaged_lines.append(row['id']) 

更新后版本

for i in range(len(coordinates)):
    for index, row in df.iterrows():
        if coordinates[i] in row['coordinates'] and sample[i]['sample_0'] == 'DS_3':
                    if row['id'] in damaged_lines:
                        break
                    else:
                        damaged_lines.append(row['id']) 
                        break

优化方案:向量化+集合匹配

核心思路是通过展开坐标列、转换为可哈希类型、利用集合快速查找,彻底避免嵌套循环,将时间复杂度从O(M*N)降至O(M+N)。

步骤1:筛选目标坐标

先从coordinates中提取样本编号为DS_3的坐标,转换为元组后存入集合(列表不可哈希,无法用于集合查找):

# 筛选DS_3对应的坐标,转成元组存入集合
target_coords = set()
for i in range(len(coordinates)):
    if sample[i]['sample_0'] == 'DS_3':
        target_coords.add(tuple(coordinates[i]))

步骤2:展开DataFrame的坐标列

将DataFrame中每条输电线路的坐标列表展开为单条记录,每个坐标对应唯一的线路ID:

import pandas as pd

# 展开坐标列,每个坐标对应一个id
expanded_df = df.explode('coordinates').reset_index(drop=True)
# 把坐标转成元组,方便和集合匹配
expanded_df['coords_tuple'] = expanded_df['coordinates'].apply(tuple)

步骤3:匹配并去重

利用isin方法快速匹配目标坐标,提取对应的线路ID后去重,得到最终结果:

# 匹配目标坐标,提取id并去重
matched_ids = expanded_df[expanded_df['coords_tuple'].isin(target_coords)]['id'].unique()
damaged_lines = list(matched_ids)

可选优化:提前过滤数据

如果有额外过滤条件(比如特定电压等级),可以在展开前先过滤DataFrame,减少后续处理的数据量:

# 示例:只保留电压为230的线路
filtered_df = df[df['voltage'] == 230]
expanded_df = filtered_df.explode('coordinates').reset_index(drop=True)

内容的提问来源于stack exchange,提问作者Kat Neumann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 17:17:04